hillclimb

What is hillclimb for?

The problem shapes it suits, the ones it does not, and the rule of thumb.

Anything you can phrase as: a program or artifact in, a number out, and a verifier that computes that number/score (potentially using data the agents never seen). Some problem types that are well suited for hillclimb:

Problem typeWhat Hillclimb improves
OptimizationPacking, routing and scheduling solutions, scored by solution quality, cost or constraint violations.
Prediction and forecastingTraining and prediction code, evaluated on hidden data using metrics such as MAE, pinball loss or CRPS.
Performance engineeringCode for a fixed workload, scored by runtime, memory use, binary size or another resource constraint.
Parameter fittingEstimation code that recovers unknown parameters, tested against cases with known ground truth.
Strategies and policiesDispatch, bidding and cache-eviction strategies, evaluated by replaying historical or simulated scenarios.
Generated artifactsSQL queries, regular expressions, solver configurations and prompts—anything that can be generated and scored.
Mathematical discoveryConstructions, counterexamples and bounds, scored by a programmatically verifiable mathematical objective.

Where it fits poorly. Verifiers that take hours: the loop needs many candidates per budget and starves. Objectives without a scalar score, like UX, prose or "nicer code". Pass/fail verifiers with no partial credit, since a 0/1 score gives the search nothing to climb. Low-dimensional continuous optimisation, where a numerical optimizer is the better tool.

Rule of thumb. Can each attempt be scored automatically in under fifteen minutes? Does the score reliably distinguish better attempts from worse ones? Does improving the score improve the outcome you actually care about? Three yeses make your problem a strong candidate for hillclimb.

How to make it hard to cheat?

Agents optimize the score your verifier produces, not necessarily the outcome you intended. If there is a shortcut or loophole, the search may find it.

Test on cases the agent never sees. Score a solver rather than a fixed answer, and evaluate it across multiple instances. Keep some instances hidden with holdout: true so the agent must discover a solution that generalizes instead of memorizing the visible cases.

Make improvements real and reproducible. Score invalid outputs as zero, avoid unnecessary rounding, and measure the verifier's noise with hillclimb verify --repeat 5. Set search.min_improvement above the noise floor so Hillclimb only accepts meaningful improvements—and inspect the winning code before using it.