problems / optional-featureshillclimb

Optional features

Unit tests, replicate metrics and reports, tunable parameters, per-instance scores.

Everything on this page is optional. A verifier that emits nothing loses nothing.

Frozen unit tests

When unit_tests is declared, hillclimb snapshots the test tree and command before any search worker in the run starts. Agents receive a disposable copy at ./unit_tests, but evaluation always uses the frozen bundle. Each parameter trial first runs the verifier, then runs the suite once; score replicates continue only after the suite passes. {python}, {solution}, and {tests} are available in the argv command, and the problem's requirements.txt must install its runner.

Candidate statuses separate correctness from executability: passing means the verifier and tests succeeded, failing means the verifier succeeded and the suite completed with failing tests, and buggy means execution crashed, timed out, or broke the evaluation contract. Failing and buggy candidates may be debugged, but only passing candidates can rank, tune, reach holdout, or ship.

Replicate metrics

Any other numeric key in the result object is journaled verbatim as the replicate's metrics ({"score": 12.3, "runtime_s": 0.8, "n_params": 40}); a trial's metrics are the per-key median of its replicates and a candidate's are its best trial's — the same rule as the score. The engine never ranks on them; they are feature dimensions for quality-diversity policies (openevolve) and context for reports. Strings, booleans and NaN are dropped silently.

Replicate reports

$HILLCLIMB_RESULT may carry a report block alongside the score: any verifier that writes one gets its breakdown stored on the replicate, rendered into improve prompts ("attack the largest contributors"), and shown by hillclimb show and the watch TUI:

{"split": "validation",
 "report": {
   "version": 1,
   "overall": {"score": 12.3, "n_origins": 100, "n_scored": 2400},
   "segment_label": "store",
   "zones": [{"zone": "store-7", "score": 19.9, "n_scored": 240}],
   "horizons": [{"bucket": "13-24h", "score": 14.1, "n": 1200}],
   "quantiles": [{"q": 0.9, "pinball": 4.1, "coverage": 0.95}],
   "worst_origins": [{"asof": "...", "zone": "store-7", "score": 44.0}],
   "residual_bias": {"mean_error": -1.2, "mean_abs_error": 8.8, "mean_actual": 41.0},
   "report_error": null
 }}

All sections are optional; order zones (any segmentation — the label is yours via segment_label) worst-first. Producers, by trust:

  • emflow problems — the evaluator computes the full breakdown (per-zone, per-horizon, per-quantile calibration, persistence skill) automatically.
  • directory problems — the verifier writes the file (see problems/circle-packing/verify.py); a verifier that discards what the solution left behind gives its report evaluator trust.
  • self-reported problems (MLE-bench) — the number is the agent's own claim, so the report is stored and rendered labelled self-reported.

Only "split": "validation" reports are ever fed back to operators — holdout evaluations never produce one, by construction. report.enabled: false in config disables prompt injection (data is still recorded).

Tunable parameters

A solution may declare its numeric knobs in params.json next to solution.py and read them through spaces.params():

params.json
{"restarts": {"type": "int", "low": 1, "high": 64, "log": true, "default": 8},
 "step":     {"type": "float", "low": 1e-4, "high": 0.1, "log": true, "default": 0.01},
 "init":     {"type": "categorical", "choices": ["grid", "random"], "default": "grid"}}
solution.py
from hillclimb import spaces
P = spaces.params()   # {"restarts": 8, "step": 0.01, "init": "grid"} — the trial's values, else the defaults

The engine then spends verifier runs, not agent turns, on that code: the search policy proposes tune actions on promising candidates, each one a new trial of the same solution.py with values from the tuner (search.tuner: random by default, optuna with pip install 'hillclimb[optuna]'), and the candidate is scored by its best trial. The greedy policy's knobs live in search.policy_params: tune_budget (extra trials per candidate, default 8, 0 disables), tune_gate (band: within the accept band of the best; best; always), tune_parallel, tune_burst. Holdout runs with the winning trial's values, best/ ships them as params.json, and a child improved from a tuned parent starts from those values as its defaults. A malformed declaration is scored on the solution's own defaults and reported in hillclimb show; seeds are never tuned — replicate variance is the noise floor, not a knob.

Per-instance scores

A third reserved key, instances, carries the breakdown of score over the problem's sub-instances — zones, folds, test cases — in the same metric and direction, with keys stable across the search:

{"score": 2.158, "instances": {"circle-00": 0.083, "circle-01": 0.083}}

They are journaled per replicate (Replicate.instance_scores), aggregated per key by median like everything else, and consumed by engines whose selection is per-instance — GEPA keeps a candidate alive if it wins on any instance, not just on average. problems/circle-packing/verify.py (one instance per circle) is the reference producer; verifiers that emit nothing lose nothing. emflow problems emit one instance per scored origin, keyed <asof>/<zone> — for GEFCom2014 that is every task x zone of the validation split (solar: 3 tasks x 3 plants = 9 instances) — so a GEPA arm on emflow://gefcom2014:solar keeps a candidate that wins any single task. A candidate that leaves an origin unscored simply lacks that key and is treated as having failed it; MLE-bench per-fold instances are a planned follow-up.