cli / commands / experimenthillclimb

hillclimb experiment

Compare setups: named arms of config overrides × repeats on a problem

hillclimb experiment COMMAND

hillclimb experiment report

Compare the arms of an experiment.

On the selected candidate's holdout score (val when holdout was off): per arm n/mean/median/spread, best-of-repeat wins, time to best and tokens; then every arm against the control, with the gap judged against the noise floor (hillclimb verify <problem> --repeat 5 measures it).

hillclimb experiment report [OPTIONS] [EXPERIMENT]

Arguments

argumentwhat it is
EXPERIMENTExperiment name or spec (default: every experiment) optional

Options

optionwhat it does
--problem TEXTFilter to one problem id.
--control TEXTArm to compare against (default: the first).
--noise-floor FLOATGap below which arms are not different (default: the spec's).
--jsonMachine-readable: the summaries as JSON (a meta-verifier reads the gaps).

hillclimb experiment run

Run an experiment: every arm × every problem × N repeats.

Sequential (the default) runs jobs in a fair order — repeat by repeat, arms round-robin inside — so shared state such as cross-search memory is seen by every arm at the same point; use it whenever an arm touches shared state. Parallel launches all jobs detached at once (machine slots still cap concurrency) — fine for stateless comparisons such as policy or model. Bounded parallel (--max-concurrent N, or the spec's max_concurrent) keeps at most N alive so no search spends its wall clock waiting for a machine slot; the launcher stays up until the last child exits, then prints the report — run it under nohup or in tmux. Unlike sequential it runs the whole matrix rather than stopping at the first failure; its exit code is 1 if any child failed, else 2 if any parked, else 0. Real agent runs — the repeat count is your cost dial.

hillclimb experiment run [OPTIONS] SPEC

Arguments

argumentwhat it is
SPECSpec YAML path, or a name under hillclimb/experiments/ required

Options

optionwhat it does
--budget TEXTPer-search budget, e.g. 10m (overrides the spec).
--repeats INTEGEROverride the spec's repeat count.
--parallel / --sequentialLaunch every search detached at once, or one after another (default: the spec's schedule).
--max-concurrent INTEGER RANGEParallel, but at most N searches alive at once: launch detached in job order, wait on the children before starting the next (overrides the spec's max_concurrent).
--run-id TEXTAppend the searches to this existing experiment run instead of creating a new one.
--first-repeat INTEGER RANGENumber the repeats from K (with --run-id: add repeats K.. to a finished run). Default: 1.
--dry-runPrint the jobs, start nothing.