Policies
What to try next — greedy, openevolve, the GEPA engine, and writing your own.
What to try next is the search engine, and it is a seam of its own: src/hillclimb/modules/policies/base.py
defines it, src/hillclimb/modules/policies/ holds the implementations. Everything else — candidate dirs,
prompts, agent calls, trials, holdout, journaling, best/ — is harness, and a policy never
touches it.
| policy | what it does |
|---|---|
greedy | debug the newest failing/buggy tip > ensemble in the final budget window > draft until num_drafts branches are scored > improve the best |
openevolve | OpenEvolve's MAP-Elites database decides what to expand: a population kept diverse over feature dimensions, split across islands with migration; parent + inspirations sampled per island (exploration / elite archive / fitness-weighted). Hillclimb's operators do the mutating, the verifier the scoring, and the debug rule is kept. pip install 'hillclimb[openevolve]' |
gepa | GEPA owns the whole loop — reflective mutation over evaluation feedback and Pareto selection over the verifier's per-instance scores — while hillclimb evaluates, journals, and holds the private holdout. A full engine, not a policy (see below). pip install 'hillclimb[gepa]' |
search:
policy: openevolve
policy_params:
num_islands: 3
feature_dimensions: [complexity, score] # built-ins: complexity, diversity, score
num_inspirations: 2 # copied in as candidate_<i>.pybest score so far per search policy, candidate by candidate. greedy is a real hillclimb run on heilbronn-convex-13; the other two curves are mock data, for the shape of the comparison rather than a measurement
Several optimizers on one problem: a mixed fleet
Repeat --policy and one run holds one search per policy, each its own engine process on the same
problem, tagged as an arm so the arms can be compared afterwards. Per-arm settings go through
--arm-set ARM:KEY=VALUE (applied after --set, which is fleet-wide); arms are named after their
policy (a repeated policy becomes greedy-2):
uv run hillclimb run circle-packing --budget 30m \
--policy greedy --policy openevolve --policy gepa \
--seed-from hillclimb/experiments/seeds/circle-packing.py \
--arm-set gepa:search.parallel_agents=1 \
--arm-set gepa:search.policy_params.max_metric_calls=60
uv run hillclimb watch # the three searches side by side, arm in the problem column
uv run hillclimb experiment report <run-id> # arms compared; --experiment NAME names it insteadThe GEPA engine is serial, so its arm needs search.parallel_agents=1 while the others keep
the fleet-wide operator count. --parallel-searches N repeats every arm N times (repeat-major,
like hillclimb experiment run). Searches of one run share live knowledge cards; pass
--set learning.enabled=false for a fair comparison, or keep it for cooperation. The same fleet
is available to embedders as hillclimb.api.run_fleet(..., engines=mixed_fleet([...])). For
repeats across problems with a committed spec, noise floors and a control arm, use
hillclimb experiment run.
Any other feature dimension must be a numeric key the verifier writes next to score (see
replicate metrics), e.g.
feature_dimensions: [runtime_s, score]. Each evolved candidate's policy_meta records its
island, grid cell and inspirations in the journal (hillclimb show <candidate> prints it); the
watch TUI and knowledge graph don't surface it yet.
The protocol
A policy is two methods over a read-only PolicyInput:
class SearchPolicy(Protocol):
name: str
params: dict # persisted into SearchMeta, so `resume` restores them
def propose(self, view: PolicyInput) -> Action | None: ...
def observe(self, view: PolicyInput, candidate: Candidate) -> None: ...propose returns one Action — an operator (draft/debug/improve/ensemble), the candidate
to target, optional inspiration_ids, and optional per-action routing — or None to hold the
slot until an in-flight result lands. observe is called after every terminal result, and
replayed over every existing candidate when the policy is constructed, which is what makes
hillclimb resume work.
Three rules the harness relies on, spelled out in modules/policies/base.py:
propose/observerun only on the scheduler thread, under the search's state lock. A policy may read candidate dirs; it must never write.- Every decision must be derivable from replayed journal state — compute it from the
PolicyInput, or rebuild your caches inobserve. - Ensemble-style actions must carry their inputs in
inspiration_ids; the harness copies those solutions into the new candidate dir.
Writing one
To add one: implement the protocol in a file and point search.policy at it, no registry edit
needed. Any value ending in .py is a policy file, relative to the folder holding the hillclimb
dir (like paths.runs_dir):
from hillclimb.modules.policies.greedy import GreedyPolicy
from hillclimb.sdk import Action
class DraftsOnly(GreedyPolicy):
name = "drafts-only"
def propose(self, view):
tip = self.debuggable_tip(view)
if tip is not None:
return Action(operator="debug", target_id=tip.candidate_id)
return self._draft_action(view)uv run hillclimb policy check --policy hillclimb/policies/drafts_only.py # before spending budget
uv run hillclimb run circle-packing --policy hillclimb/policies/drafts_only.py
uv run hillclimb run circle-packing --policy greedy --policy hillclimb/policies/drafts_only.py # fleet: arm "drafts_only"The file exposes its policy as the one class it defines with propose and observe, or as
POLICY = <class or factory>; the constructor gets params (from search.policy_params) and
complexity_start when it accepts them. search.yaml records the path as written and
policy_sha256, the file's hash at search start: the identity of an edited exploration process,
the way seed_sha256 identifies a seed. resume reloads the file from the same path and warns
when the hash changed, since replay may then diverge. Built-in names (greedy, openevolve) stay
in the _POLICIES dict in modules/policies/__init__.py. modules/policies/greedy.py is under 300 lines and is
the reference. Beam search, MCTS, evolutionary populations, novelty search and bandits over
operators all fit this shape — greedy is just the one that ships.
The whole exploration process is one dict plus one file. Every knob greedy reads — num_drafts,
max_debug_depth, the ensemble* window and the tune_* budget — comes from
search.policy_params first and falls back to the config block it historically lived in, so an
existing config.yaml behaves as before and an experiment arm (or, later, an agent editing the
loop) is handed a single dict; GreedyPolicy.resolved_params(config) is that dict fully resolved.
Before spending an agent hour on an edited process, run the conformance check:
uv run hillclimb policy check [--policy NAME] [--set search.policy_params.k=v] [--problem P] [--smoke]It replays every recorded journal in the store (plus an empty one) through the policy with no
agent or verifier and reports each contract breach it can see: a hold with empty slots on an empty
journal (the search would never start), two fresh instances disagreeing at some budget point
(resume would diverge), a target or inspiration id that does not exist, an operator that needs a
target without one, a debug on a non-failing/non-buggy candidate, a mutated journal or a file
written under a search dir, a factory that hands back the same object, and a prompt override that
lints dirty. --smoke --problem P then runs a short --agent dummy search so the whole loop,
prompts included, executes once; --json is the machine-readable form. Exit 1 on any breach.
Operator prompts are part of that process too. Any template in hillclimb/prompts/<name>.md
(paths.prompts_dir) shadows the package template of the same name (draft, improve, debug,
ensemble, the contract_* templates, the cue snippets); an override may drop {{tokens}} but
never add one the engine does not fill — the engine refuses to start on such a file. Every search
records templates_sha256 and templates_overridden in search.yaml, so two searches are
comparable only when their prompt hashes agree.
The GEPA engine (optional extra)
Some optimizers cannot be reduced to "what next?" — they own proposal, reflection, and selection
themselves. Those integrate one tier up, as a search strategy
(src/hillclimb/harness/glue.py): dispatched by the same search.policy name, but handed the
full dependency set instead of a PolicyInput. The architecture is docs/optimizer-host-plan.md;
GEPA is the first such engine:
- greedy: hillclimb chooses the parent and asks an agent to mutate;
openevolvepolicy: OpenEvolve supplies selection/inspirations, hillclimb's agent still mutates;gepaengine: GEPA drives reflective mutation and Pareto search, hillclimb evaluates and records.
uv sync --extra gepa
uv run hillclimb run <problem> --policy gepa --seed-from my_solution.pyGEPA's mutations are performed by a routed hillclimb agent (configure routing.gepa, falling back
to routing.default and the global agent/model) in scratch dirs under
SEARCH_DIR/gepa/proposals/; every evaluation becomes a normal journaled cNNN candidate, so
watch, tree, and chart work unchanged (policy_meta.optimizer == "gepa" carries the
lineage). GEPA checkpoints under SEARCH_DIR/gepa/state and hillclimb resume continues both the
journal and the optimizer, with a warm evaluation cache so replayed proposals cost nothing.
MVP limits: one mutable file (solution.py), serial (search.parallel_agents: 1), no merge,
and a required executable seed — pass --seed-from or ship an executable baseline. Holdout
privacy is strict and one-way: holdout scoring runs only after the optimizer finishes, and no
holdout value ever reaches GEPA's prompts, feedback, or state (regression-tested with sentinels).
When the verifier emits per-instance scores, frontier_type: instance (the default) tracks GEPA's
Pareto frontier per instance; without them the frontier degenerates to the aggregate score.