Agents
Who writes the code — claude-code, codex, pi, dummy — and how operators are routed.
Operators are headless coding-agent processes, one per operator call, behind the agent seam in
src/hillclimb/agents/:
| agent | what it is |
|---|---|
claude-code | Claude Code in headless mode — the production agent; bills your Claude subscription |
codex | Codex CLI in non-interactive mode; uses your Codex login by default |
codex + agent_auth: openrouter | the same Codex CLI pointed at OpenRouter: cheap open models billed to OpenRouter credits, no subscription touched |
pi | pi coding agent in JSON mode; subscription login, API keys, OpenRouter or custom local providers, with per-operator sampling |
dummy | no model calls: a scripted operator for exercising the engine, TUIs and run layout |
fake | deterministic canned operator for the test suite |
Other agents (OpenCode, …) plug in at the same seam: an agent implements the Agent
protocol in agents/base.py — take a prompt plus a working directory, return the agent's JSON
result — and is selected with --agent <name>.
Every agent runs confined: it writes only to its candidate's folder and cannot read your keys. See Sandbox.
Cheap operators through OpenRouter
agent: codex
agent_auth: openrouter
model: qwen/qwen3-coder # any OpenRouter model idOPENROUTER_API_KEY comes from the environment or a .env beside config.yaml. Routing mixes
providers per operator, and a models: pool lets the bandit learn which cheap model actually
earns improvements:
routing:
draft: {agent: claude-code, agent_auth: subscription, model: sonnet}
improve: {models: [qwen/qwen3-coder, deepseek/deepseek-v3]}
debug: {model: cohere/north-mini-code:free}Codex resends an identical 12k-token preamble on every call, so prefer models whose providers
cache prompts: the journaled $0.53 at qwen3-coder prices) and
another spent its whole agent timeout without converging. Running out of credits parks the search
— top up, then cache_read_input_tokens tells you whether the discount is landing.
Set budget.max_cost_usd — cheap per token is not cheap per search, because a weaker model
compensates with volume: one measured DRAFT burned 3M tokens (resume. To compare models head to head, give an experiment one arm per model
(arm_overrides: {model: …}); the chart and experiment report group on the arm tags.
Sampling with pi
Install pi and select a provider-qualified model. Subscription mode copies the credentials from
pi's own /login (~/.pi/agent/auth.json); api-key uses provider environment variables such as
ANTHROPIC_API_KEY. OpenRouter requires OPENROUTER_API_KEY in the environment or .env beside
config.yaml:
agent: pi
agent_auth: openrouter
model: openrouter/deepseek/deepseek-v3.2
routing:
draft: {sampling: {temperature: 1.0, top_p: 0.95}}
improve: {sampling: {temperature: 0.2}}| setting | meaning |
|---|---|
routing.<op>.sampling | Numeric provider request parameters, e.g. temperature, top_p, top_k, min_p; only supported by pi |
routing.default.sampling | Fallback for all operators; an operator's dict replaces it, and {} disables inherited sampling |
pi.models_file | Optional pi models.json for custom providers, including vLLM and llama.cpp |
Sampling follows action → operator → default routing precedence and is recorded on each candidate.
Unsupported agent combinations fail config validation. A short, tool-free preflight checks each
distinct pi model and sampling combination (including every model in a pool) before search work
starts. Provider rejection fails startup, including errors pi emits with exit code 0. Preflight
streams and accounting are under pi-preflight/; their small provider cost is separate from
candidate spend.
Pi runs with an isolated PI_CODING_AGENT_DIR under ~/.cache/hillclimb/pi-home/<auth>/, with
discovery of personal extensions, skills, prompt templates and context files disabled. Custom
model files get a content-hashed subdirectory so concurrent searches cannot overwrite each other's
providers. PI_OFFLINE=1 disables startup updates and telemetry; install/update pi explicitly to
update its bundled catalogue. Tested with pi 0.73.1. Its update --models flag is not supported.
Debug children fork the parent's pi session into their own candidate directory. Raw pi events
remain in agent_stream.jsonl and are visible in watch, including token usage and pi's reported
cost. For OpenRouter models with zero reported cost, Hillclimb falls back to its pricing
catalogue.
For a local OpenAI-compatible server, set agent_auth: api-key,
model: vllm/Qwen/Qwen3-Coder-30B-A3B-Instruct and point pi.models_file at a file like:
{"providers": {"vllm": {"baseUrl": "http://localhost:8000/v1",
"api": "openai-completions", "apiKey": "none",
"models": [{"id": "Qwen/Qwen3-Coder-30B-A3B-Instruct", "contextWindow": 131072}]}}}Relative model-file paths resolve from the project root; with an explicit config-file load they resolve beside that config. Provider/model support determines which sampling fields are accepted. Use the preflight to check compatibility; reasoning effort remains a separate follow-up.
hillclimb/experiments/temperature.yaml compares temperatures 0.2, 0.7 and 1.0 on
circle-packing, with three repeats and three concurrent searches. Run
hillclimb experiment run temperature --dry-run to inspect the jobs;
hillclimb experiment report temperature --json reports the arm verdicts after the experiment
finishes.
Operator scaffolds and model routing
Two prompt scaffolds sharpen the default operators (both on by default; the operators: config
block gates prompt injection only, so A/B arms record identical data):
- Retrieval-augmented draft (
operators.draft_retrieval) — the draft agent is told to web-search the current state of the art for the problem class before writing code (methods only — searching for solutions to the specific competition is explicitly forbidden). - Ablation-guided improve (
operators.improve_ablation) — the improve agent first attributes the score to the solution's components (fast, subsampled ablation runs, focused by the trial report's breakdown), records findings inablation.md, then confines its ONE change to the highest-leverage component. Later improves of the same solution are handed the newest siblingablation.mdso components aren't re-measured.
The routing: block maps operators to agents/models; giving a route a models: pool instead
of a scalar turns model choice into a UCB1 bandit (per operator) that learns which model earns
improvements — rewards derive from journaled results (improved on parent = 1, working-but-flat =
0.25, buggy = 0), so bandit state rebuilds from journal replay and survives resume:
routing:
improve: {models: [sonnet, opus-4.8]} # bandit picks per call
debug: {model: haiku} # scalar routes stay scalars