Experiments
Which setup wins? Arms, repeats, and a verdict judged against the noise floor.
Every knob — search policy, its params, the model, cross-search memory, trial count — is a config
setting, so "does X help?" is one experiment: a problem × named arms (sets of config overrides)
× N repeats. A spec lives in hillclimb/experiments/<name>.yaml:
problems: [circle-packing]
repeats: 3
budget: 15m
schedule: sequential # sequential | parallel
max_concurrent: 8 # parallel only: searches alive at once
noise_floor: 0.02 # from `hillclimb verify circle-packing --repeat 5`
arms:
greedy: {search.policy: greedy} # first arm = the control
greedy-nomem: {search.policy: greedy, learning.enabled: false}
openevolve: {search.policy: openevolve, search.policy_params: {population_size: 50}}hillclimb experiment run <name> launches the matrix: sequentially by default — repeat by repeat,
arms round-robin inside, so shared state such as the knowledge graph is seen by every arm at the
same point (mandatory when an arm touches memory) — or --parallel for stateless comparisons
(policy, model). Parallel without a bound starts every search at once, and past the machine's
agent slots (concurrency.machine_max_agents, default min(8, cores-2)) the rest burn their
wall clock in waiting-slot — so give it max_concurrent: N in the spec (or
--max-concurrent N): the launcher starts searches in job order, waits on its children before
starting the next, runs the whole matrix, and prints the report at the end (run it under nohup
or in tmux; exit 1 if any child failed, 2 if any parked). To add repeats to a finished run,
--run-id <run> --first-repeat K --repeats M appends repeats K..K+M-1 into the same run, so report
and chart keep grouping as one experiment. Each search is a normal
hillclimb run … --experiment <name> --arm <arm> --set key=value, tagged in its search.yaml
(experiment, arm, repeat, arm_overrides), so a search you start by hand with those flags
counts too. hillclimb experiment report <name> compares the arms on the selected candidate's
holdout score (val when holdout is off): n / mean / median / spread, best-of-repeat wins, minutes
to best, tokens, and each arm's paired gap to the control judged against the noise floor — a gap
inside it is reported as "within noise, not a result". hillclimb chart colours an experiment's
curves by arm.
A spec may name one shared executable seed — seed_from: seeds/foo.py, resolved against the
spec's directory — which rides --seed-from into every child search, so arms are compared from
identical source (the dry run prints the resolved path and its sha256). Mandatory for engines that
require a seed (GEPA); see hillclimb/experiments/gepa-vs-openevolve-vs-greedy.yaml for the
three-strategy comparison this shipped with, and
gepa-vs-openevolve-vs-greedy-heilbronn.yaml for the same three engines across a difficulty
ladder (problems/heilbronn-{11,14,17}, stamped by problems/make_heilbronn.py; its seed reads N
off the problem, so one seed_from serves every level). One hillclimb chart per problem (a bare
hillclimb chart lists them, p cycles them, c in hillclimb watch opens the highlighted one)
and hillclimb similarity <run> for the arm-coloured map (similarity reference <run> for the
cube).
every search of one problem, as hillclimb chart plots it: each dot is a scored candidate in landing order, the line is the best score so far, and the published scores from problem.yaml are the reference lines