Budget and parallelism
A gate rather than a wall, and the two levels of concurrency.
The budget is wall-clock time: hillclimb keeps drafting and scoring candidates until it is
spent, and that is the only stopping rule.
The budget is a gate, not a wall
A search stops starting operators once its budget is inside the stop margin; by default
(budget.deadline: graceful) whatever is still in flight finishes and is committed, so a search
can overrun by up to one operator. The duration column in hillclimb watch keeps counting and
says by how much: 1h 04m 16s (budget: 1h, 4m 16s over). Pass --set budget.deadline=hard (or
set it in config.yaml) to cut in-flight operators off at the deadline instead; they are
journaled abandoned ("cut off at the budget deadline").
Two levels of parallelism
Parallelism in hillclimb has two levels: --parallel-searches is how many independent searches
(exploration trees) attack the problem (each its own engine process, all in one run so they share
what they learn), --parallel-agents how many coding agents each search keeps busy at once,
each generating one candidate at a time running in the background.
hillclimb run heilbronn-convex-13 --parallel-searches 2 --parallel-agents 3Concurrency is bounded machine-wide, not per search: concurrency.parallel_agents is how many
agents one search keeps in flight, and concurrency.machine_max_agents (default
min(8, cores - 2), 0 = off) caps the total across every search on the machine — extra
agents wait (waiting-slot in hillclimb watch). Every verifier and agent process gets
OMP/OPENBLAS/MKL_NUM_THREADS=1 unless the parent environment sets them, so N agents cost at
most N cores; hillclimb ps shows what is actually running.
Cost
Set budget.max_cost_usd — cheap per token is not cheap per search, because a weaker model
compensates with volume.