Why build it
Different models are good at different work, and choosing between them by feel does not scale past one person or one repository. I wanted the choice to be a rule I could inspect, the execution to leave evidence I could review, and the definition of "done" to stay with me rather than with whichever model finished last.
What it does
| Cross-LLM division | Claude and GPT-series models, driven through their existing CLIs, take on planning, focused analysis, implementation, and independent review as separate roles. |
|---|---|
| Deterministic routing | The implementing model is chosen by explicit task type. General implementation and concurrency-sensitive work take different paths, so quality and cost considerations are rules rather than habits. |
| Reusable executor | Initialization, environment discovery, an isolated worktree bound to a specific commit, staged execution, and a run record for every task. |
| Acceptance state | A successful model exit still needs verification. Inputs, commits, and run evidence are saved; quality approval is never granted automatically. |
| Applied in practice | A JobSignal repair pilot and the LLM Interview Lab delivery. Routing and run checks were adjusted after real failures. |
The flow, and who does what
- Task and acceptance criteria me
- Planning and focused analysis coordinated
- Choose the implementing model by task type CLI
- Execute in an isolated workspace, save run evidence CLI
- Tests and independent review coordinated
- Fix, or adjust the task routing coordinated
- Human acceptance and release me
CLI = implemented as a stage in the local executor · coordinated = driven by a main session with the CLI and human judgment · me = my decision
A real failure and what changed
On one task the implementation stage exited successfully, but the saved run evidence did not cover the acceptance criteria. Nothing in the process treated that exit as done: the task moved to awaiting verification, the gap was identified in review, and the work was redone. The routing rule and the run checks were tightened afterwards so the same gap would be caught earlier. This is the behavior the whole design exists for.
My contribution
- Defined the model responsibilities and the acceptance boundary between execution success and task completion.
- Built the reusable executor: initialization, environment discovery, commit-bound isolated worktrees, staged execution, and run records.
- Revised routing and run checks based on real failures rather than on a diagram.
- Remained responsible for the final delivery on every task that went through it.
Status and boundaries
Local, installable CLI · version 0.1 Single-stage executor Automatic scheduling not implemented
The CLI is a private preview package with a single-stage executor. The full flow is coordinated by a main session; automatic scheduling, persistent state, cross-process coordination, automatic review admission, and hard budget limits are future scope. I do not claim a measured cost reduction: the only figures I have are price scenarios under identical token and caching assumptions, not a controlled comparison.
Related public work
agent-harness-pack is the public, earlier half of this work: role templates for implementation and independent review, delivery rules, CI templates, and lessons learned. It is a separate repository from the CLI and is not the CLI's source code.