Documentation
contributehillclimb

Contribute

Three seams, each a small protocol. Implement one and it plugs in.

hillclimb has three seams, and each one is a small protocol: how to search, who writes the code, and where the records live. Implement one and it plugs in.

the engine and its seams
You or your agent call the hillclimb CLI, which runs the search, scores every candidate through the verifier, journals it and keeps the best. It executes each step by running one headless coding agent such as Claude Code or Codex. you, or your top-level agent hillclimb CLI runs the search, scores through the verifier, journals every candidate, keeps the best claude code codex … headless coding agents as operators

simple and detailed are the same diagram at two levels of detail — the buttons above the drawing switch between them

01 — A search policy

Decides what to try next: one action at a time over a read-only view of the search. Everything else (prompts, agent calls, trials, journaling) stays with the harness.

src/hillclimb/modules/policies/base.py
class SearchPolicy(Protocol):
    name: str
    params: dict

    # one Action -- draft | debug | improve | ensemble --
    # or None to hold the slot open
    def propose(self, view: SearchView) -> Action | None: ...

    # every terminal result, and replayed on resume
    def observe(self, view: SearchView, candidate: Candidate) -> None: ...

greedy is the only one so far, at 170 lines. Beam search, MCTS, evolutionary populations or a bandit over operators all fit this shape. Add yours to the registry in modules/policies/__init__.py and run it with hillclimb run --policy yours.

02 — An agent

Runs one headless coding agent over a prompt and a working directory, and reports what happened.

src/hillclimb/agents/base.py
class Agent(Protocol):
    name: str

    # take a prompt and a working directory, write code, report what happened
    def invoke(self, request: OperatorRequest) -> OperatorResult: ...

claude-code and codex are supported today; dummy runs the engine with no model calls at all. Future integration could include OpenCode, Pi and friends.

03 — A data store

Holds a search's records: metadata, the append-only candidate journal, the status heartbeat, and the stop/prune queue. Candidate code and agent logs stay on disk either way.

src/hillclimb/harness/store.py
class DataStore(Protocol):
    # upserts for runs, searches and status; appends for the journal;
    # a consume-once queue for commands
    def record_run(self, meta: RunMeta) -> None: ...
    def record_search(self, meta: SearchMeta) -> None: ...
    def runs(self) -> list[RunMeta]: ...
    def searches(self, problem_key=None, run_id=None) -> list[SearchRecord]: ...
    def search(self, key: SearchKey) -> SearchRecord | None: ...
    def journal(self, key: SearchKey) -> JournalBackend: ...
    def write_status(self, key: SearchKey, status: SearchStatus) -> None: ...
    def read_status(self, key: SearchKey) -> SearchStatus | None: ...
    def enqueue_command(self, key: SearchKey, cmd: ControlCommand) -> None: ...
    def drain_commands(self, key: SearchKey) -> list[ControlCommand]: ...
    def clear_stale_stops(self, key: SearchKey) -> None: ...
    def close(self) -> None: ...

files is the zero-setup default; sqlite (store.backend: sqlite) puts the same records in one database file, safe for many engines writing at once.


Suggest new policies, agents and stores as pull requests: the registries are plain dicts and the store is a config key. The shortest complete examples are greedy.py, dummy.py and store.py's FileDataStore.

Licence

hillclimb is free (free both as in freedom and "free beer") software under the MIT License.