clihillclimb

CLI reference

Every command, and what it does.

Search-addressing commands take <run-id>/<search-id>, a bare <run-id> (when the run has a single search), or latest (the default).

Every flag, every default

This page is the short view: which command you want, and why. All commands has one page per command — arguments, options and defaults — generated from the CLI itself.

Problems and runs

commandwhat it does
problem get [circle-packing|knapsack|heilbronn-convex-13]copy a bundled problem into hillclimb/problems/ (creates the hillclimb/ dir if needed) and list its files (fetch is a deprecated alias)
demo [--budget 10m] [--parallel-searches 3] [--parallel-agents 3]problem get + run --parallel-searches in one command
init [dir]create the hillclimb/ dir (config, problems/, specs/, runs/) with an example problem
verify <problem> [--repeat N] [--holdout]run a problem's verifier once, outside a search; --repeat measures the noise floor
run <target> [--name ...] [--budget 2h] [--agent ...] [--model ...]start a run for one problem or a suite YAML
run <problem> --parallel-searches N --parallel-agents MN independent searches (detached engines, one run) each running M agents at once
run <problem> --set key=value … [--experiment E --arm A]any config setting, dotted; tag the search as an experiment arm
resume [search]continue a parked / stopped / crashed search
smoke [problem]one real agent call end-to-end (auth / contract check)

Watching a climb

commandwhat it does
status [search]search state + candidate tree (text)
watchlive TUI over runs, searches, and candidates
chartlive chart: best score so far by tested-candidate count across the problem's searches as a staircase, every scored candidate a dot (one line per arm in an experiment)
tree [search]live 3D exploration tree of one search: colour is the operator, silhouette the fate (expanded / discontinued / best / failed); j/k scrub through time
tree2 [search]the same tree drawn like the Darwin Gödel Machine's archive: the candidate number inside each circle, fill = score (viridis, bright = best; hollow = no working solution), ring = what the search did with it (white expanded — the spine the policy walked / none a scored, scored and left / red failed), star = best, bold white path = the best's lineage; circles are sized to the zoom so they never overlap, numbers appear as they grow; the legend toggles each stage, the best and the lineage (1-5 or click)
archive [search]the tree2 archive tree on the left and the progress chart on the right — every scored candidate at (candidate number, score), the best-so-far staircase, and the lineage of the final best as a thick line, the same parent chain drawn bold in the tree; j/k scrub both panels together, click a node to ring its dot on the chart; the chart's legend sits in the corner the climb leaves empty and toggles its series (6-9 or click), the hover readout keeps off it
surface [search]live 3D fitness surface: the search's candidates on the problem's terrain (needs a landscape.py in the problem; problems/fitness-landscape/ is the reference)
similarity map [search] [--single] [--metric M]live 3D map: every candidate embedded by pairwise distance (behavioral by default; m cycles structural and blend), so nearby dots are alike — lineage edges, a gold best-so-far trail, hover reads distances, click dims everything outside a lineage, space replays the search growing; an experiment arm opens its whole run, coloured by arm; a bare similarity is this view
similarity reference [search] [--single]live 3D cube: each candidate at behavioral / structural / lineage distance from the search's seed (or baseline; c toggles the champion); an experiment arm opens its whole run, coloured by arm, n/p stepping through the run's problems; a problem's fingerprint.py defines the behavioral axis; v swaps between the two views
graphthe knowledge-graph TUI (same screen as knowledge graph)
show [search] <candidate-id>everything about one candidate: scores, evaluation breakdown, diff vs parent, output
hillclimb archive
archive tree
progress

what hillclimb archive draws: the archive tree beside the progress chart, scrubbed together

Controlling a climb

commandwhat it does
psevery process hillclimb owns on this machine: engines with their agents and verifiers nested; orphan marks engines whose hillclimb dir was deleted
stop [search] [--all]graceful stop: finish current operator, then park; --all also reaps orphaned engines when no hillclimb dir is found
kill [search]SIGTERM the engine now (state finalized, resumable)
prune <search> <candidate-id>cut a candidate and its subtree from the search
summit [problem]copy the best solution found so far across every run of a problem next to your hillclimb/ folder; works mid-climb

Memory

commandwhat it does
knowledge graph [--stats]interactive knowledge-graph TUI (or a text summary)
knowledge rebuildforce-rebuild the derived knowledge/graph.json index
knowledge distill [search] [--backfill]run the LLM claims pass on a search / all cards
knowledge backfilldistill cards from every finished search that lacks one
knowledge live [run]the live cards concurrent searches in a run are sharing
paper add <pdf> [--problem <target>] / paper listdistill a PDF into knowledge claims that seed future searches (one agent pass per paper, content-hash cached)
knowledge consolidate [--dry-run]sleep phase: generalize claims + rewrite playbooks
knowledge query "<terms>" [--json]read-only memory lookup (also available to agents)
knowledge show <target>the prior-experience section a new search would get

Experiments and policies

commandwhat it does
experiment run <spec> [--repeats N] [--budget B] [--parallel] [--max-concurrent N] [--run-id R --first-repeat K] [--dry-run]every arm × problem × repeat of a spec; --max-concurrent bounds how many run at once, --run-id appends repeats to a finished run
experiment report [spec] [--problem X] [--control A] [--noise-floor F] [--json]compare the arms on holdout; --json gives a meta-verifier the gaps and verdicts as data
policy check [--policy NAME] [--set k=v] [--problem P] [--smoke] [--json]conformance check for a search policy over the store's recorded journals; --smoke adds a dummy-agent search

Exit code 2 from run/resume means the search parked or was stopped — resume it.