Benchmark problems
emflow, MLE-bench and Einstein Arena targets, and the provider registry behind them.
Targets of the form <scheme>://<name> resolve through a lazy BenchmarkProvider registry
instead of adding target-specific branches to the runner.
emflow problems
With the emflow extra installed (pip install 'hillclimb[emflow]'), targets of the form
emflow://<name> run problems from emflow's registry —
agents author Predictor classes (solution.py exposing get_model()), and the provider
supplies the verifier command: a generic evaluator fits and scores them on the problem's
validation split, with the hidden holdout as a second run of the same command. A bare package name
is a virtual suite (one search per variant):
uv run hillclimb run emflow://gefcom2014:solar --budget 2h # one track
uv run hillclimb run emflow://gefcom2014 --budget 2h # all four tracksThe baseline candidate (c000) is the benchmark's reference model evaluated for real, and a
finished search ends with one official emflow Verifier run (leaderboard row + rank, with
n_trials recorded for selection honesty). Programmatic use:
hillclimb.run_search("emflow://gefcom2014:solar", budget_s=7200).
MLE-bench problems
Targets of the form mlebench://<competition-id> run
MLE-bench competitions against a local mle-bench checkout
(located via paths.mlebench_python; prepare data first with mlebench prepare -c <competition-id> in that venv). Agents see only the prepared PUBLIC split and climb on their own
validation score; when the search finishes, the engine runs mlebench grade-sample exactly once
on the selected candidate and writes the report (score + medal flags) to mlebench-grade.json —
the private test set never influences selection.
A split name is a virtual suite, one search per listed competition (lite is an alias for the
22-competition low split):
uv run hillclimb run mlebench://spaceship-titanic --budget 2h # one competition
uv run hillclimb run mlebench://lite --budget 4h # MLE-bench LiteEinstein Arena problems
einsteinarena://<slug> resolves a public Einstein Arena
construction problem into a normal verifier-backed ProblemSpec. Hillclimb fetches only the
public problem and leaderboard endpoints, hashes the fields that define evaluation, and runs the
downloaded evaluate(data) -> float verifier locally. Candidates write submission.json;
Hillclimb never registers an agent, submits a solution, downloads an incumbent, or posts to a
discussion.
uv run hillclimb run einsteinarena://circle-packing --budget 10m
uv run hillclimb run einsteinarena://smoke --budget 10m # three-problem pilot suiteThe first resolution caches a content-addressed snapshot under
~/.cache/hillclimb/benchmark-problems/einsteinarena/. search.yaml records the pinned
@sha256:<revision> target, so resume is reproducible and can run offline. The verifier source is
public but untrusted code: it runs in the managed local runtime with credentials scrubbed, not in
Einstein Arena's E2B sandbox. Override the API root for a mirror or test deployment with:
einsteinarena:
base_url: https://einsteinarena.com
request_timeout_s: 30Adding a provider
Benchmark integrations use the lazy BenchmarkProvider registry rather than adding
target-specific branches to the runner. A provider implements load_problem() and
resolve_target() (plus optional chart baselines), then registers a URI scheme with
hillclimb.register_benchmark_provider(...). Search engines consume the resulting ProblemSpec
unchanged.