hillclimb
CLI reference
Every command, and what it does.
Search-addressing commands take <run-id>/<search-id>, a bare <run-id> (when the run has a
single search), or latest (the default).
Every flag, every default
This page is the short view: which command you want, and why. All commands has one page per command — arguments, options and defaults — generated from the CLI itself.
Problems and runs
| command | what it does |
|---|---|
problem get [circle-packing|knapsack|heilbronn-convex-13] | copy a bundled problem into hillclimb/problems/ (creates the hillclimb/ dir if needed) and list its files (fetch is a deprecated alias) |
demo [--budget 10m] [--parallel-searches 3] [--parallel-operators 3] | problem get + run --parallel-searches in one command |
init [dir] | create the hillclimb/ dir (config, problems/, specs/, runs/) with an example problem |
verify <problem> [--repeat N] [--holdout] | run a problem's verifier once, outside a search; --repeat measures the noise floor |
run <target> [--name ...] [--budget 2h] [--backend ...] [--model ...] | start a run for one problem or a suite YAML |
run <problem> --parallel-searches N --parallel-operators M | N independent searches (detached engines, one run) each running M agents at once |
run <problem> --set key=value … [--experiment E --arm A] | any config setting, dotted; tag the search as an experiment arm |
resume [search] | continue a parked / stopped / crashed search |
smoke [problem] | one real agent call end-to-end (auth / contract check) |
Watching a climb
| command | what it does |
|---|---|
status [search] | search state + candidate tree (text) |
watch | live TUI over runs, searches, and candidates |
chart | live chart: best score so far by tested-candidate count across the problem's searches as a staircase, every scored candidate a dot (one line per arm in an experiment) |
tree [search] | live 3D exploration tree of one search: colour is the operator, silhouette the fate (expanded / discontinued / best / failed); j/k scrub through time |
tree2 [search] | the same tree drawn like the Darwin Gödel Machine's archive: the candidate number inside each circle, fill = score (viridis, bright = best; hollow = no working solution), ring = what the search did with it (white expanded — the spine the policy walked / none a scored, scored and left / red failed), star = best, bold white path = the best's lineage; circles are sized to the zoom so they never overlap, numbers appear as they grow; the legend toggles each stage, the best and the lineage (1-5 or click) |
archive [search] | the tree2 archive tree on the left and the progress chart on the right — every scored candidate at (candidate number, score), the best-so-far staircase, and the lineage of the final best as a thick line, the same parent chain drawn bold in the tree; j/k scrub both panels together, click a node to ring its dot on the chart; the chart's legend sits in the corner the climb leaves empty and toggles its series (6-9 or click), the hover readout keeps off it |
surface [search] | live 3D fitness surface: the search's candidates on the problem's terrain (needs a landscape.py in the problem; problems/fitness-landscape/ is the reference) |
similarity map [search] [--single] [--metric M] | live 3D map: every candidate embedded by pairwise distance (behavioral by default; m cycles structural and blend), so nearby dots are alike — lineage edges, a gold best-so-far trail, hover reads distances, click dims everything outside a lineage, space replays the search growing; an experiment arm opens its whole run, coloured by arm; a bare similarity is this view |
similarity reference [search] [--single] | live 3D cube: each candidate at behavioral / structural / lineage distance from the search's seed (or baseline; c toggles the champion); an experiment arm opens its whole run, coloured by arm, n/p stepping through the run's problems; a problem's fingerprint.py defines the behavioral axis; v swaps between the two views |
graph | the knowledge-graph TUI (same screen as knowledge graph) |
show [search] <candidate-id> | everything about one candidate: scores, evaluation breakdown, diff vs parent, output |
hillclimb archive
archive tree
progress
what hillclimb archive draws: the archive tree beside the progress chart, scrubbed together
Controlling a climb
| command | what it does |
|---|---|
ps | every process hillclimb owns on this machine: engines with their agents and verifiers nested; orphan marks engines whose hillclimb dir was deleted |
stop [search] [--all] | graceful stop: finish current operator, then park; --all also reaps orphaned engines when no hillclimb dir is found |
kill [search] | SIGTERM the engine now (state finalized, resumable) |
prune <search> <candidate-id> | cut a candidate and its subtree from the search |
summit [problem] | copy the best solution found so far across every run of a problem next to your hillclimb/ folder; works mid-climb |
Memory
| command | what it does |
|---|---|
knowledge graph [--stats] | interactive knowledge-graph TUI (or a text summary) |
knowledge rebuild | force-rebuild the derived knowledge/graph.json index |
knowledge distill [search] [--backfill] | run the LLM claims pass on a search / all cards |
knowledge backfill | distill cards from every finished search that lacks one |
knowledge live [run] | the live cards concurrent searches in a run are sharing |
paper add <pdf> [--problem <target>] / paper list | distill a PDF into knowledge claims that seed future searches (one agent pass per paper, content-hash cached) |
knowledge consolidate [--dry-run] | sleep phase: generalize claims + rewrite playbooks |
knowledge query "<terms>" [--json] | read-only memory lookup (also available to agents) |
knowledge show <target> | the prior-experience section a new search would get |
Experiments and policies
| command | what it does |
|---|---|
experiment run <spec> [--repeats N] [--budget B] [--parallel] [--max-concurrent N] [--run-id R --first-repeat K] [--dry-run] | every arm × problem × repeat of a spec; --max-concurrent bounds how many run at once, --run-id appends repeats to a finished run |
experiment report [spec] [--problem X] [--control A] [--noise-floor F] [--json] | compare the arms on holdout; --json gives a meta-verifier the gaps and verdicts as data |
policy check [--policy NAME] [--set k=v] [--problem P] [--smoke] [--json] | conformance check for a search policy over the store's recorded journals; --smoke adds a dummy-backend search |
Exit code 2 from run/resume means the search parked or was stopped — resume it.