anatomy / experimentshillclimb

Experiments

Which setup wins? Arms, repeats, and a verdict judged against the noise floor.

Every knob — search policy, its params, the model, cross-search memory, trial count — is a config setting, so "does X help?" is one experiment: a problem × named arms (sets of config overrides) × N repeats. A spec lives in hillclimb/experiments/<name>.yaml:

hillclimb/experiments/policies.yaml
problems: [circle-packing]
repeats: 3
budget: 15m
schedule: sequential        # sequential | parallel
max_concurrent: 8           # parallel only: searches alive at once
noise_floor: 0.02           # from `hillclimb verify circle-packing --repeat 5`
arms:
  greedy:       {search.policy: greedy}             # first arm = the control
  greedy-nomem: {search.policy: greedy, learning.enabled: false}
  openevolve:   {search.policy: openevolve, search.policy_params: {population_size: 50}}

hillclimb experiment run <name> launches the matrix: sequentially by default — repeat by repeat, arms round-robin inside, so shared state such as the knowledge graph is seen by every arm at the same point (mandatory when an arm touches memory) — or --parallel for stateless comparisons (policy, model). Parallel without a bound starts every search at once, and past the machine's agent slots (concurrency.machine_max_agents, default min(8, cores-2)) the rest burn their wall clock in waiting-slot — so give it max_concurrent: N in the spec (or --max-concurrent N): the launcher starts searches in job order, waits on its children before starting the next, runs the whole matrix, and prints the report at the end (run it under nohup or in tmux; exit 1 if any child failed, 2 if any parked). To add repeats to a finished run, --run-id <run> --first-repeat K --repeats M appends repeats K..K+M-1 into the same run, so report and chart keep grouping as one experiment. Each search is a normal hillclimb run … --experiment <name> --arm <arm> --set key=value, tagged in its search.yaml (experiment, arm, repeat, arm_overrides), so a search you start by hand with those flags counts too. hillclimb experiment report <name> compares the arms on the selected candidate's holdout score (val when holdout is off): n / mean / median / spread, best-of-repeat wins, minutes to best, tokens, and each arm's paired gap to the control judged against the noise floor — a gap inside it is reported as "within noise, not a result". hillclimb chart colours an experiment's curves by arm.

A spec may name one shared executable seed — seed_from: seeds/foo.py, resolved against the spec's directory — which rides --seed-from into every child search, so arms are compared from identical source (the dry run prints the resolved path and its sha256). Mandatory for engines that require a seed (GEPA); see hillclimb/experiments/gepa-vs-openevolve-vs-greedy.yaml for the three-strategy comparison this shipped with, and gepa-vs-openevolve-vs-greedy-heilbronn.yaml for the same three engines across a difficulty ladder (problems/heilbronn-{11,14,17}, stamped by problems/make_heilbronn.py; its seed reads N off the problem, so one seed_from serves every level). One hillclimb chart per problem (a bare hillclimb chart lists them, p cycles them, c in hillclimb watch opens the highlighted one) and hillclimb similarity <run> for the arm-coloured map (similarity reference <run> for the cube).

hillclimb chart

every search of one problem, as hillclimb chart plots it: each dot is a scored candidate in landing order, the line is the best score so far, and the published scores from problem.yaml are the reference lines