hillclimb experiment
Compare setups: named arms of config overrides × repeats on a problem
hillclimb experiment COMMANDhillclimb experiment report
Compare the arms of an experiment.
On the selected candidate's holdout score (val when holdout was off): per arm n/mean/median/spread, best-of-repeat wins, time to best and tokens; then every arm against the control, with the gap judged against the noise floor (hillclimb verify <problem> --repeat 5 measures it).
hillclimb experiment report [OPTIONS] [EXPERIMENT]Arguments
| argument | what it is |
|---|---|
EXPERIMENT | Experiment name or spec (default: every experiment) optional |
Options
| option | what it does |
|---|---|
--problem TEXT | Filter to one problem id. |
--control TEXT | Arm to compare against (default: the first). |
--noise-floor FLOAT | Gap below which arms are not different (default: the spec's). |
--json | Machine-readable: the summaries as JSON (a meta-verifier reads the gaps). |
hillclimb experiment run
Run an experiment: every arm × every problem × N repeats.
Sequential (the default) runs jobs in a fair order — repeat by repeat, arms round-robin inside — so shared state such as cross-search memory is seen by every arm at the same point; use it whenever an arm touches shared state. Parallel launches all jobs detached at once (machine slots still cap concurrency) — fine for stateless comparisons such as policy or model. Bounded parallel (--max-concurrent N, or the spec's max_concurrent) keeps at most N alive so no search spends its wall clock waiting for a machine slot; the launcher stays up until the last child exits, then prints the report — run it under nohup or in tmux. Unlike sequential it runs the whole matrix rather than stopping at the first failure; its exit code is 1 if any child failed, else 2 if any parked, else 0. Real agent runs — the repeat count is your cost dial.
hillclimb experiment run [OPTIONS] SPECArguments
| argument | what it is |
|---|---|
SPEC | Spec YAML path, or a name under hillclimb/experiments/ required |
Options
| option | what it does |
|---|---|
--budget TEXT | Per-search budget, e.g. 10m (overrides the spec). |
--repeats INTEGER | Override the spec's repeat count. |
--parallel / --sequential | Launch every search detached at once, or one after another (default: the spec's schedule). |
--max-concurrent INTEGER RANGE | Parallel, but at most N searches alive at once: launch detached in job order, wait on the children before starting the next (overrides the spec's max_concurrent). |
--run-id TEXT | Append the searches to this existing experiment run instead of creating a new one. |
--first-repeat INTEGER RANGE | Number the repeats from K (with --run-id: add repeats K.. to a finished run). Default: 1. |
--dry-run | Print the jobs, start nothing. |