hillclimb

Use Claude Code and Codex to autonomously optimize a score you define.

You describe one problem as a single script: verifier.sh runs a candidate solution.py and writes back a number. hillclimb then spends a compute budget having headless coding agents write, debug and improve that solution.py — scoring every version through your verifier and keeping the one that scores best.

the climb

hillclimb chart
best score so far per search policy, candidate by candidate. greedy is a real hillclimb run on heilbronn-convex-13; the other two curves are mock data, for the shape of the comparison rather than a measurement