anatomy / agentshillclimb

Agents

Who writes the code — claude-code, codex, pi, dummy — and how operators are routed.

Operators are headless coding-agent processes, one per operator call, behind the agent seam in src/hillclimb/agents/:

agentwhat it is
claude-codeClaude Code in headless mode — the production agent; bills your Claude subscription
codexCodex CLI in non-interactive mode; uses your Codex login by default
codex + agent_auth: openrouterthe same Codex CLI pointed at OpenRouter: cheap open models billed to OpenRouter credits, no subscription touched
pipi coding agent in JSON mode; subscription login, API keys, OpenRouter or custom local providers, with per-operator sampling
dummyno model calls: a scripted operator for exercising the engine, TUIs and run layout
fakedeterministic canned operator for the test suite

Other agents (OpenCode, …) plug in at the same seam: an agent implements the Agent protocol in agents/base.py — take a prompt plus a working directory, return the agent's JSON result — and is selected with --agent <name>.

Every agent runs confined: it writes only to its candidate's folder and cannot read your keys. See Sandbox.

Cheap operators through OpenRouter

hillclimb/config.yaml
agent: codex
agent_auth: openrouter
model: qwen/qwen3-coder          # any OpenRouter model id

OPENROUTER_API_KEY comes from the environment or a .env beside config.yaml. Routing mixes providers per operator, and a models: pool lets the bandit learn which cheap model actually earns improvements:

routing:
  draft:   {agent: claude-code, agent_auth: subscription, model: sonnet}
  improve: {models: [qwen/qwen3-coder, deepseek/deepseek-v3]}
  debug:   {model: cohere/north-mini-code:free}

Codex resends an identical 12k-token preamble on every call, so prefer models whose providers cache prompts: the journaled cache_read_input_tokens tells you whether the discount is landing. Set budget.max_cost_usd — cheap per token is not cheap per search, because a weaker model compensates with volume: one measured DRAFT burned 3M tokens ($0.53 at qwen3-coder prices) and another spent its whole agent timeout without converging. Running out of credits parks the search — top up, then resume. To compare models head to head, give an experiment one arm per model (arm_overrides: {model: …}); the chart and experiment report group on the arm tags.

Sampling with pi

Install pi and select a provider-qualified model. Subscription mode copies the credentials from pi's own /login (~/.pi/agent/auth.json); api-key uses provider environment variables such as ANTHROPIC_API_KEY. OpenRouter requires OPENROUTER_API_KEY in the environment or .env beside config.yaml:

agent: pi
agent_auth: openrouter
model: openrouter/deepseek/deepseek-v3.2
routing:
  draft:   {sampling: {temperature: 1.0, top_p: 0.95}}
  improve: {sampling: {temperature: 0.2}}
settingmeaning
routing.<op>.samplingNumeric provider request parameters, e.g. temperature, top_p, top_k, min_p; only supported by pi
routing.default.samplingFallback for all operators; an operator's dict replaces it, and {} disables inherited sampling
pi.models_fileOptional pi models.json for custom providers, including vLLM and llama.cpp

Sampling follows action → operator → default routing precedence and is recorded on each candidate. Unsupported agent combinations fail config validation. A short, tool-free preflight checks each distinct pi model and sampling combination (including every model in a pool) before search work starts. Provider rejection fails startup, including errors pi emits with exit code 0. Preflight streams and accounting are under pi-preflight/; their small provider cost is separate from candidate spend.

Pi runs with an isolated PI_CODING_AGENT_DIR under ~/.cache/hillclimb/pi-home/<auth>/, with discovery of personal extensions, skills, prompt templates and context files disabled. Custom model files get a content-hashed subdirectory so concurrent searches cannot overwrite each other's providers. PI_OFFLINE=1 disables startup updates and telemetry; install/update pi explicitly to update its bundled catalogue. Tested with pi 0.73.1. Its update --models flag is not supported.

Debug children fork the parent's pi session into their own candidate directory. Raw pi events remain in agent_stream.jsonl and are visible in watch, including token usage and pi's reported cost. For OpenRouter models with zero reported cost, Hillclimb falls back to its pricing catalogue.

For a local OpenAI-compatible server, set agent_auth: api-key, model: vllm/Qwen/Qwen3-Coder-30B-A3B-Instruct and point pi.models_file at a file like:

{"providers": {"vllm": {"baseUrl": "http://localhost:8000/v1",
  "api": "openai-completions", "apiKey": "none",
  "models": [{"id": "Qwen/Qwen3-Coder-30B-A3B-Instruct", "contextWindow": 131072}]}}}

Relative model-file paths resolve from the project root; with an explicit config-file load they resolve beside that config. Provider/model support determines which sampling fields are accepted. Use the preflight to check compatibility; reasoning effort remains a separate follow-up.

hillclimb/experiments/temperature.yaml compares temperatures 0.2, 0.7 and 1.0 on circle-packing, with three repeats and three concurrent searches. Run hillclimb experiment run temperature --dry-run to inspect the jobs; hillclimb experiment report temperature --json reports the arm verdicts after the experiment finishes.

Operator scaffolds and model routing

Two prompt scaffolds sharpen the default operators (both on by default; the operators: config block gates prompt injection only, so A/B arms record identical data):

  • Retrieval-augmented draft (operators.draft_retrieval) — the draft agent is told to web-search the current state of the art for the problem class before writing code (methods only — searching for solutions to the specific competition is explicitly forbidden).
  • Ablation-guided improve (operators.improve_ablation) — the improve agent first attributes the score to the solution's components (fast, subsampled ablation runs, focused by the trial report's breakdown), records findings in ablation.md, then confines its ONE change to the highest-leverage component. Later improves of the same solution are handed the newest sibling ablation.md so components aren't re-measured.

The routing: block maps operators to agents/models; giving a route a models: pool instead of a scalar turns model choice into a UCB1 bandit (per operator) that learns which model earns improvements — rewards derive from journaled results (improved on parent = 1, working-but-flat = 0.25, buggy = 0), so bandit state rebuilds from journal replay and survives resume:

routing:
  improve: {models: [sonnet, opus-4.8]}   # bandit picks per call
  debug: {model: haiku}                   # scalar routes stay scalars