Optional features
Unit tests, replicate metrics and reports, tunable parameters, per-instance scores.
Everything on this page is optional. A verifier that emits nothing loses nothing.
Frozen unit tests
When unit_tests is declared, hillclimb snapshots the test tree and command before any search
worker in the run starts. Agents receive a disposable copy at ./unit_tests, but evaluation
always uses the frozen bundle. Each parameter trial first runs the verifier, then runs the suite
once; score replicates continue only after the suite passes. {python}, {solution}, and
{tests} are available in the argv command, and the problem's requirements.txt must install its
runner.
Candidate statuses separate correctness from executability: passing means the verifier and tests
succeeded, failing means the verifier succeeded and the suite completed with failing tests, and
buggy means execution crashed, timed out, or broke the evaluation contract. Failing and buggy
candidates may be debugged, but only passing candidates can rank, tune, reach holdout, or ship.
Replicate metrics
Any other numeric key in the result object is journaled verbatim as the replicate's metrics
({"score": 12.3, "runtime_s": 0.8, "n_params": 40}); a trial's metrics are the per-key median of
its replicates and a candidate's are its best trial's — the same rule as the score. The engine
never ranks on them; they are feature dimensions for quality-diversity policies (openevolve) and
context for reports. Strings, booleans and NaN are dropped silently.
Replicate reports
$HILLCLIMB_RESULT may carry a report block alongside the score: any verifier that writes one
gets its breakdown stored on the replicate, rendered into improve prompts ("attack the largest
contributors"), and shown by hillclimb show and the watch TUI:
{"split": "validation",
"report": {
"version": 1,
"overall": {"score": 12.3, "n_origins": 100, "n_scored": 2400},
"segment_label": "store",
"zones": [{"zone": "store-7", "score": 19.9, "n_scored": 240}],
"horizons": [{"bucket": "13-24h", "score": 14.1, "n": 1200}],
"quantiles": [{"q": 0.9, "pinball": 4.1, "coverage": 0.95}],
"worst_origins": [{"asof": "...", "zone": "store-7", "score": 44.0}],
"residual_bias": {"mean_error": -1.2, "mean_abs_error": 8.8, "mean_actual": 41.0},
"report_error": null
}}All sections are optional; order zones (any segmentation — the label is yours via
segment_label) worst-first. Producers, by trust:
- emflow problems — the evaluator computes the full breakdown (per-zone, per-horizon, per-quantile calibration, persistence skill) automatically.
- directory problems — the verifier writes the file (see
problems/circle-packing/verify.py); a verifier that discards what the solution left behind gives its report evaluator trust. - self-reported problems (MLE-bench) — the number is the agent's own claim, so the report is stored and rendered labelled self-reported.
Only "split": "validation" reports are ever fed back to operators — holdout evaluations never
produce one, by construction. report.enabled: false in config disables prompt injection (data is
still recorded).
Tunable parameters
A solution may declare its numeric knobs in params.json next to solution.py and read them
through spaces.params():
{"restarts": {"type": "int", "low": 1, "high": 64, "log": true, "default": 8},
"step": {"type": "float", "low": 1e-4, "high": 0.1, "log": true, "default": 0.01},
"init": {"type": "categorical", "choices": ["grid", "random"], "default": "grid"}}from hillclimb import spaces
P = spaces.params() # {"restarts": 8, "step": 0.01, "init": "grid"} — the trial's values, else the defaultsThe engine then spends verifier runs, not agent turns, on that code: the search policy proposes
tune actions on promising candidates, each one a new trial of the same solution.py with
values from the tuner (search.tuner: random by default, optuna with
pip install 'hillclimb[optuna]'), and the candidate is scored by its best trial. The greedy
policy's knobs live in search.policy_params: tune_budget (extra trials per candidate, default
8, 0 disables), tune_gate (band: within the accept band of the best; best; always),
tune_parallel, tune_burst. Holdout runs with the winning trial's values, best/ ships them as
params.json, and a child improved from a tuned parent starts from those values as its defaults.
A malformed declaration is scored on the solution's own defaults and reported in hillclimb show;
seeds are never tuned — replicate variance is the noise floor, not a knob.
Per-instance scores
A third reserved key, instances, carries the breakdown of score over the problem's
sub-instances — zones, folds, test cases — in the same metric and direction, with keys stable
across the search:
{"score": 2.158, "instances": {"circle-00": 0.083, "circle-01": 0.083}}They are journaled per replicate (Replicate.instance_scores), aggregated per key by median like
everything else, and consumed by engines whose selection is per-instance — GEPA keeps a candidate
alive if it wins on any instance, not just on average. problems/circle-packing/verify.py (one
instance per circle) is the reference producer; verifiers that emit nothing lose nothing. emflow
problems emit one instance per scored origin, keyed <asof>/<zone> — for GEFCom2014 that is every
task x zone of the validation split (solar: 3 tasks x 3 plants = 9 instances) — so a GEPA arm on
emflow://gefcom2014:solar keeps a candidate that wins any single task. A candidate that leaves an
origin unscored simply lacks that key and is treated as having failed it; MLE-bench per-fold
instances are a planned follow-up.