What is hillclimb for?
The problem shapes it suits, the ones it does not, and the rule of thumb.
Anything you can phrase as: a program or artifact in, a number out, and a verifier that computes
that number/score (potentially using data the agents never seen). Some problem types that are well
suited for hillclimb:
| Problem type | What Hillclimb improves |
|---|---|
| Optimization | Packing, routing and scheduling solutions, scored by solution quality, cost or constraint violations. |
| Prediction and forecasting | Training and prediction code, evaluated on hidden data using metrics such as MAE, pinball loss or CRPS. |
| Performance engineering | Code for a fixed workload, scored by runtime, memory use, binary size or another resource constraint. |
| Parameter fitting | Estimation code that recovers unknown parameters, tested against cases with known ground truth. |
| Strategies and policies | Dispatch, bidding and cache-eviction strategies, evaluated by replaying historical or simulated scenarios. |
| Generated artifacts | SQL queries, regular expressions, solver configurations and prompts—anything that can be generated and scored. |
| Mathematical discovery | Constructions, counterexamples and bounds, scored by a programmatically verifiable mathematical objective. |
Where it fits poorly. Verifiers that take hours: the loop needs many candidates per budget and starves. Objectives without a scalar score, like UX, prose or "nicer code". Pass/fail verifiers with no partial credit, since a 0/1 score gives the search nothing to climb. Low-dimensional continuous optimisation, where a numerical optimizer is the better tool.
Rule of thumb. Can each attempt be scored automatically in under fifteen minutes? Does the
score reliably distinguish better attempts from worse ones? Does improving the score improve the
outcome you actually care about? Three yeses make your problem a strong candidate for hillclimb.
How to make it hard to cheat?
Agents optimize the score your verifier produces, not necessarily the outcome you intended. If there is a shortcut or loophole, the search may find it.
Test on cases the agent never sees. Score a solver rather than a fixed answer, and evaluate it
across multiple instances. Keep some instances hidden with holdout: true so the agent must
discover a solution that generalizes instead of memorizing the visible cases.
Make improvements real and reproducible. Score invalid outputs as zero, avoid unnecessary
rounding, and measure the verifier's noise with hillclimb verify --repeat 5. Set
search.min_improvement above the noise floor so Hillclimb only accepts meaningful
improvements—and inspect the winning code before using it.