Introduction
Auto-hillclimbing for verifier-defined problems.
You describe one problem as a single script: verifier.sh runs a candidate solution.py and
writes back a number. hillclimb then spends a compute budget having headless coding agents
write, debug and improve that solution.py — scoring every version through your verifier and
keeping the one that scores best.
best score so far per search policy, candidate by candidate. greedy is a real hillclimb run on heilbronn-convex-13; the other two curves are mock data, for the shape of the comparison rather than a measurement
A greedy search engine spawns headless coding agents (Claude Code) as operators — draft new
solutions, debug failures, improve the best one, ensemble at the end — executes every
candidate in a sandboxed venv, scores it against a hidden holdout, and keeps the best submission
in runs/<run-id>/searches/<search-id>/best/.
Three layers
The design is deliberately three-layered:
hillclimb CLI — the headless engine. Scriptable, plain output, meaningful exit codes. All
run state lives on disk.
Your interactive agent (Claude Code) is the front door: the repo ships a skill
(.claude/skills/hillclimb/SKILL.md) that teaches it to start, monitor, and control runs in the
background while you chat.
hillclimb watch — a live TUI you keep open beside the agent: runs → searches → candidate
trees, with an on-demand candidate detail panel for notes, scores, lineage, output, and the
timestamped operator stream when present. A candidate still in flight gets a live console under
its overview: the agent's stream, then the verifier's stdout/stderr, appended as they are written
(tail -f style — it follows the end until you scroll up, and f follows again). Drag the
divider, use + / - to resize the detail panel, or m to maximize it.
Where to start
Quickstart
Install the CLI, get a bundled problem, start climbing.
Run, search, candidate
How a climb is put together, coarse to fine.
What is hillclimb for?
The problem shapes it suits, and the ones it does not.
Walkthrough
The Heilbronn triangle problem end to end, with the run's own charts.
Defining problems
A problem is a folder, and a problem is its verifier.