indexhillclimb

Introduction

Auto-hillclimbing for verifier-defined problems.

You describe one problem as a single script: verifier.sh runs a candidate solution.py and writes back a number. hillclimb then spends a compute budget having headless coding agents write, debug and improve that solution.py — scoring every version through your verifier and keeping the one that scores best.

hillclimb chart

best score so far per search policy, candidate by candidate. greedy is a real hillclimb run on heilbronn-convex-13; the other two curves are mock data, for the shape of the comparison rather than a measurement

A greedy search engine spawns headless coding agents (Claude Code) as operators — draft new solutions, debug failures, improve the best one, ensemble at the end — executes every candidate in a sandboxed venv, scores it against a hidden holdout, and keeps the best submission in runs/<run-id>/searches/<search-id>/best/.

Three layers

The design is deliberately three-layered:

hillclimb CLI — the headless engine. Scriptable, plain output, meaningful exit codes. All run state lives on disk.

Your interactive agent (Claude Code) is the front door: the repo ships a skill (.claude/skills/hillclimb/SKILL.md) that teaches it to start, monitor, and control runs in the background while you chat.

hillclimb watch — a live TUI you keep open beside the agent: runs → searches → candidate trees, with an on-demand candidate detail panel for notes, scores, lineage, output, and the timestamped operator stream when present. A candidate still in flight gets a live console under its overview: the agent's stream, then the verifier's stdout/stderr, appended as they are written (tail -f style — it follows the end until you scroll up, and f follows again). Drag the divider, use + / - to resize the detail panel, or m to maximize it.

Where to start