Verifier contract
Exit 0 means valid; the score is a file, not a line of stdout.
The verifier is the scoring process the engine starts. It drives solution.py itself — run it,
import it, shell out to it — and reports the score:
#!/usr/bin/env bash
set -euo pipefail
"$HILLCLIMB_PYTHON" "$HILLCLIMB_SOLUTION" # writes ./submission.csv
rm -f "$HILLCLIMB_RESULT" # only the scorer may score
exec "$HILLCLIMB_PYTHON" problem/verify.py # writes $HILLCLIMB_RESULT| exit 0 | the verifier accepted the candidate; declared unit tests must still pass |
| non-zero | the candidate is buggy and routes to the debug operator |
$HILLCLIMB_RESULT | the score: {"score": <float>, "report": {...}, <other numeric keys>}, or a bare number |
$HILLCLIMB_PYTHON | the managed runtime venv's interpreter (bare python resolves via PATH: wrong interpreter) |
$HILLCLIMB_SOLUTION | the solution path for this run (replicate-dir aware) |
$HILLCLIMB_SPLIT | validation or holdout |
$HILLCLIMB_REPLICATE_SEED | set when the engine runs repeated replicates (also exported as the legacy $HILLCLIMB_TRIAL_SEED) |
The result file is both the score carrier and the completion proof: the engine deletes it before every run, so a stale file can never fake success, and exit 0 without one is a contract violation rather than a silent zero. Reading the score from a file rather than stdout is what keeps it honest — agent-authored code runs inside the verifier and shares its stdout.
The command runs with cwd = the candidate's working dir (solution.py, plus ./problem/ and
./data/ symlinks). Validation runs get a credential-scrubbed environment; --holdout runs in a
directory agents never see, with the full environment (private holdout data may be gated).
Every run is confined by the sandbox: it writes only to its candidate's folder,
and it has no network unless the problem sets allow_internet_during_solution: true.
uv run hillclimb verify my-problem --repeat 5 # score it outside a searchhillclimb verify is the fastest way to check a new verifier, and --repeat reports the spread
between identical runs — an improvement smaller than that is noise, not progress.
problems/bin-packing/ and problems/circle-packing/ are the two reference shapes
(evaluator-driven, and run-then-score).
Make it hard to fool
Whatever the verifier fails to check, the search will eventually exploit. Score invalid outputs as
zero, score a solver rather than a fixed answer, and keep some instances hidden with
holdout: true.