anatomy / search-policieshillclimb

Policies

What to try next — greedy, openevolve, the GEPA engine, and writing your own.

What to try next is the search engine, and it is a seam of its own: src/hillclimb/modules/policies/base.py defines it, src/hillclimb/modules/policies/ holds the implementations. Everything else — candidate dirs, prompts, agent calls, trials, holdout, journaling, best/ — is harness, and a policy never touches it.

policywhat it does
greedydebug the newest failing/buggy tip > ensemble in the final budget window > draft until num_drafts branches are scored > improve the best
openevolveOpenEvolve's MAP-Elites database decides what to expand: a population kept diverse over feature dimensions, split across islands with migration; parent + inspirations sampled per island (exploration / elite archive / fitness-weighted). Hillclimb's operators do the mutating, the verifier the scoring, and the debug rule is kept. pip install 'hillclimb[openevolve]'
gepaGEPA owns the whole loop — reflective mutation over evaluation feedback and Pareto selection over the verifier's per-instance scores — while hillclimb evaluates, journals, and holds the private holdout. A full engine, not a policy (see below). pip install 'hillclimb[gepa]'
config.yaml — OpenEvolve's quality-diversity search over hillclimb's operators
search:
  policy: openevolve
  policy_params:
    num_islands: 3
    feature_dimensions: [complexity, score]   # built-ins: complexity, diversity, score
    num_inspirations: 2                       # copied in as candidate_<i>.py
hillclimb chart

best score so far per search policy, candidate by candidate. greedy is a real hillclimb run on heilbronn-convex-13; the other two curves are mock data, for the shape of the comparison rather than a measurement

Several optimizers on one problem: a mixed fleet

Repeat --policy and one run holds one search per policy, each its own engine process on the same problem, tagged as an arm so the arms can be compared afterwards. Per-arm settings go through --arm-set ARM:KEY=VALUE (applied after --set, which is fleet-wide); arms are named after their policy (a repeated policy becomes greedy-2):

uv run hillclimb run circle-packing --budget 30m \
  --policy greedy --policy openevolve --policy gepa \
  --seed-from hillclimb/experiments/seeds/circle-packing.py \
  --arm-set gepa:search.parallel_agents=1 \
  --arm-set gepa:search.policy_params.max_metric_calls=60
uv run hillclimb watch                          # the three searches side by side, arm in the problem column
uv run hillclimb experiment report <run-id>     # arms compared; --experiment NAME names it instead

The GEPA engine is serial, so its arm needs search.parallel_agents=1 while the others keep the fleet-wide operator count. --parallel-searches N repeats every arm N times (repeat-major, like hillclimb experiment run). Searches of one run share live knowledge cards; pass --set learning.enabled=false for a fair comparison, or keep it for cooperation. The same fleet is available to embedders as hillclimb.api.run_fleet(..., engines=mixed_fleet([...])). For repeats across problems with a committed spec, noise floors and a control arm, use hillclimb experiment run.

Any other feature dimension must be a numeric key the verifier writes next to score (see replicate metrics), e.g. feature_dimensions: [runtime_s, score]. Each evolved candidate's policy_meta records its island, grid cell and inspirations in the journal (hillclimb show <candidate> prints it); the watch TUI and knowledge graph don't surface it yet.

The protocol

A policy is two methods over a read-only PolicyInput:

src/hillclimb/modules/policies/base.py
class SearchPolicy(Protocol):
    name: str
    params: dict   # persisted into SearchMeta, so `resume` restores them

    def propose(self, view: PolicyInput) -> Action | None: ...
    def observe(self, view: PolicyInput, candidate: Candidate) -> None: ...

propose returns one Action — an operator (draft/debug/improve/ensemble), the candidate to target, optional inspiration_ids, and optional per-action routing — or None to hold the slot until an in-flight result lands. observe is called after every terminal result, and replayed over every existing candidate when the policy is constructed, which is what makes hillclimb resume work.

Three rules the harness relies on, spelled out in modules/policies/base.py:

  • propose/observe run only on the scheduler thread, under the search's state lock. A policy may read candidate dirs; it must never write.
  • Every decision must be derivable from replayed journal state — compute it from the PolicyInput, or rebuild your caches in observe.
  • Ensemble-style actions must carry their inputs in inspiration_ids; the harness copies those solutions into the new candidate dir.

Writing one

To add one: implement the protocol in a file and point search.policy at it, no registry edit needed. Any value ending in .py is a policy file, relative to the folder holding the hillclimb dir (like paths.runs_dir):

hillclimb/policies/drafts_only.py
from hillclimb.modules.policies.greedy import GreedyPolicy
from hillclimb.sdk import Action


class DraftsOnly(GreedyPolicy):
    name = "drafts-only"

    def propose(self, view):
        tip = self.debuggable_tip(view)
        if tip is not None:
            return Action(operator="debug", target_id=tip.candidate_id)
        return self._draft_action(view)
uv run hillclimb policy check --policy hillclimb/policies/drafts_only.py   # before spending budget
uv run hillclimb run circle-packing --policy hillclimb/policies/drafts_only.py
uv run hillclimb run circle-packing --policy greedy --policy hillclimb/policies/drafts_only.py  # fleet: arm "drafts_only"

The file exposes its policy as the one class it defines with propose and observe, or as POLICY = <class or factory>; the constructor gets params (from search.policy_params) and complexity_start when it accepts them. search.yaml records the path as written and policy_sha256, the file's hash at search start: the identity of an edited exploration process, the way seed_sha256 identifies a seed. resume reloads the file from the same path and warns when the hash changed, since replay may then diverge. Built-in names (greedy, openevolve) stay in the _POLICIES dict in modules/policies/__init__.py. modules/policies/greedy.py is under 300 lines and is the reference. Beam search, MCTS, evolutionary populations, novelty search and bandits over operators all fit this shape — greedy is just the one that ships.

The whole exploration process is one dict plus one file. Every knob greedy reads — num_drafts, max_debug_depth, the ensemble* window and the tune_* budget — comes from search.policy_params first and falls back to the config block it historically lived in, so an existing config.yaml behaves as before and an experiment arm (or, later, an agent editing the loop) is handed a single dict; GreedyPolicy.resolved_params(config) is that dict fully resolved. Before spending an agent hour on an edited process, run the conformance check:

uv run hillclimb policy check [--policy NAME] [--set search.policy_params.k=v] [--problem P] [--smoke]

It replays every recorded journal in the store (plus an empty one) through the policy with no agent or verifier and reports each contract breach it can see: a hold with empty slots on an empty journal (the search would never start), two fresh instances disagreeing at some budget point (resume would diverge), a target or inspiration id that does not exist, an operator that needs a target without one, a debug on a non-failing/non-buggy candidate, a mutated journal or a file written under a search dir, a factory that hands back the same object, and a prompt override that lints dirty. --smoke --problem P then runs a short --agent dummy search so the whole loop, prompts included, executes once; --json is the machine-readable form. Exit 1 on any breach.

Operator prompts are part of that process too. Any template in hillclimb/prompts/<name>.md (paths.prompts_dir) shadows the package template of the same name (draft, improve, debug, ensemble, the contract_* templates, the cue snippets); an override may drop {{tokens}} but never add one the engine does not fill — the engine refuses to start on such a file. Every search records templates_sha256 and templates_overridden in search.yaml, so two searches are comparable only when their prompt hashes agree.

The GEPA engine (optional extra)

Some optimizers cannot be reduced to "what next?" — they own proposal, reflection, and selection themselves. Those integrate one tier up, as a search strategy (src/hillclimb/harness/glue.py): dispatched by the same search.policy name, but handed the full dependency set instead of a PolicyInput. The architecture is docs/optimizer-host-plan.md; GEPA is the first such engine:

  • greedy: hillclimb chooses the parent and asks an agent to mutate;
  • openevolve policy: OpenEvolve supplies selection/inspirations, hillclimb's agent still mutates;
  • gepa engine: GEPA drives reflective mutation and Pareto search, hillclimb evaluates and records.
uv sync --extra gepa
uv run hillclimb run <problem> --policy gepa --seed-from my_solution.py

GEPA's mutations are performed by a routed hillclimb agent (configure routing.gepa, falling back to routing.default and the global agent/model) in scratch dirs under SEARCH_DIR/gepa/proposals/; every evaluation becomes a normal journaled cNNN candidate, so watch, tree, and chart work unchanged (policy_meta.optimizer == "gepa" carries the lineage). GEPA checkpoints under SEARCH_DIR/gepa/state and hillclimb resume continues both the journal and the optimizer, with a warm evaluation cache so replayed proposals cost nothing.

MVP limits: one mutable file (solution.py), serial (search.parallel_agents: 1), no merge, and a required executable seed — pass --seed-from or ship an executable baseline. Holdout privacy is strict and one-way: holdout scoring runs only after the optimizer finishes, and no holdout value ever reaches GEPA's prompts, feedback, or state (regression-tested with sentinels). When the verifier emits per-instance scores, frontier_type: instance (the default) tracks GEPA's Pareto frontier per instance; without them the frontier degenerates to the aggregate score.