cli / run-layouthillclimb

Runs on disk

How runs and searches are laid out, and the DataStore behind them.

runs/
└── <run-id>/                        # one hillclimb invocation
    ├── run.yaml                     # run metadata (schema_version: 2)
    ├── logs/                        # per-search engine logs (suite runs)
    └── searches/<search-id>/        # <problem-id>, then <problem-id>-2, -3 for more on one problem
        ├── search.yaml              # immutable search config (schema_version: 2)
        ├── status.json              # live heartbeat: state, pid, budget, current candidate
        ├── journal.jsonl            # append-only event log — the source of truth
        ├── control/                 # command queue (stop/prune) polled by the engine
        ├── best/                    # current selected submission (+ solution.py)
        └── candidates/<candidate-id>/  # one working dir per operator call
            ├── prompt.md
            ├── agent_stream.jsonl
            ├── solution.py
            └── submission.csv

Single-writer rule

Only the engine process mutates search state. The TUI, the CLI control commands, and chat agents all send commands through the store's command queue (control/ in the file backend), or apply them offline only when the engine is provably not running. Never edit journal.jsonl or status.json by hand.

The DataStore (store.backend)

run.yaml, search.yaml, journal.jsonl, status.json and control/ are the file backend's representation of a search's records. The engine, the CLI and the TUIs all read and write those records through one abstraction — the DataStore (src/hillclimb/harness/store.py) — and hillclimb/config.yaml picks the backend:

hillclimb/config.yaml
store:
  backend: files        # default — the folder above; nothing to set up, git-versionable
  # backend: sqlite     # one database file instead: hillclimb/store.sqlite
  # sqlite_path: hillclimb/store.sqlite

With sqlite, a search dir holds only what has to be files (candidates/, best/, logs) and everything else lives in the database — cross-run views (the chart, store searches, experiments) query it instead of walking run dirs, and N concurrent engines (the demo) write it safely. The single-writer rule is unchanged: the engine owns a search's records whichever backend holds them; stop/prune go through the store's command queue.

hillclimb store sync imports the folder's searches into the configured store (skipping ones it already has) — run it once after switching to sqlite so earlier history shows up. hillclimb store searches [--problem KEY] lists what the store holds.

A new backend implements the DataStore protocol: run/search metadata (upsert), the journal (append-only, returned in append order — policies replay it), one status record per search, and a consume-once command queue. Candidate working dirs, best/, agent streams/logs, problems, knowledge YAML and agent slots stay on the local filesystem in every backend — agents and verifiers need real files.

Search states: running (fresh heartbeat + live pid) · parked (rate limit; resume later) · stopped (user stop/kill) · done · failed (see last_error) · crashed (derived: stale heartbeat or dead pid) · unknown (no status.json yet).