problems / verifier-contracthillclimb

Verifier contract

Exit 0 means valid; the score is a file, not a line of stdout.

The verifier is the scoring process the engine starts. It drives solution.py itself — run it, import it, shell out to it — and reports the score:

problems/my-problem/verifier.sh
#!/usr/bin/env bash
set -euo pipefail

"$HILLCLIMB_PYTHON" "$HILLCLIMB_SOLUTION"   # writes ./submission.csv

rm -f "$HILLCLIMB_RESULT"                   # only the scorer may score
exec "$HILLCLIMB_PYTHON" problem/verify.py  # writes $HILLCLIMB_RESULT
exit 0the verifier accepted the candidate; declared unit tests must still pass
non-zerothe candidate is buggy and routes to the debug operator
$HILLCLIMB_RESULTthe score: {"score": <float>, "report": {...}, <other numeric keys>}, or a bare number
$HILLCLIMB_PYTHONthe managed runtime venv's interpreter (bare python resolves via PATH: wrong interpreter)
$HILLCLIMB_SOLUTIONthe solution path for this run (replicate-dir aware)
$HILLCLIMB_SPLITvalidation or holdout
$HILLCLIMB_REPLICATE_SEEDset when the engine runs repeated replicates (also exported as the legacy $HILLCLIMB_TRIAL_SEED)

The result file is both the score carrier and the completion proof: the engine deletes it before every run, so a stale file can never fake success, and exit 0 without one is a contract violation rather than a silent zero. Reading the score from a file rather than stdout is what keeps it honest — agent-authored code runs inside the verifier and shares its stdout.

The command runs with cwd = the candidate's working dir (solution.py, plus ./problem/ and ./data/ symlinks). Validation runs get a credential-scrubbed environment; --holdout runs in a directory agents never see, with the full environment (private holdout data may be gated).

Every run is confined by the sandbox: it writes only to its candidate's folder, and it has no network unless the problem sets allow_internet_during_solution: true.

uv run hillclimb verify my-problem --repeat 5   # score it outside a search

hillclimb verify is the fastest way to check a new verifier, and --repeat reports the spread between identical runs — an improvement smaller than that is noise, not progress. problems/bin-packing/ and problems/circle-packing/ are the two reference shapes (evaluator-driven, and run-then-score).

Make it hard to fool

Whatever the verifier fails to check, the search will eventually exploit. Score invalid outputs as zero, score a solver rather than a fixed answer, and keep some instances hidden with holdout: true.