← All work

Developer tooling · Cross-LLM Development Workflow

Task routing, independent review, and verifiable delivery.

A reusable local CLI and a coordinated process for splitting development work across LLM providers, running each task in an isolated workspace with saved run evidence, and keeping "the model exited successfully" separate from "the task is accepted".

Local CLI preview Applied in project workflows Private source

Why build it

Different models are good at different work, and choosing between them by feel does not scale past one person or one repository. I wanted the choice to be a rule I could inspect, the execution to leave evidence I could review, and the definition of "done" to stay with me rather than with whichever model finished last.

What it does

Cross-LLM divisionClaude and GPT-series models, driven through their existing CLIs, take on planning, focused analysis, implementation, and independent review as separate roles.
Deterministic routingThe implementing model is chosen by explicit task type. General implementation and concurrency-sensitive work take different paths, so quality and cost considerations are rules rather than habits.
Reusable executorInitialization, environment discovery, an isolated worktree bound to a specific commit, staged execution, and a run record for every task.
Acceptance stateA successful model exit still needs verification. Inputs, commits, and run evidence are saved; quality approval is never granted automatically.
Applied in practiceA JobSignal repair pilot and the LLM Interview Lab delivery. Routing and run checks were adjusted after real failures.

The flow, and who does what

  1. Task and acceptance criteria me
  2. Planning and focused analysis coordinated
  3. Choose the implementing model by task type CLI
  4. Execute in an isolated workspace, save run evidence CLI
  5. Tests and independent review coordinated
  6. Fix, or adjust the task routing coordinated
  7. Human acceptance and release me

CLI = implemented as a stage in the local executor · coordinated = driven by a main session with the CLI and human judgment · me = my decision

A real failure and what changed

On one task the implementation stage exited successfully, but the saved run evidence did not cover the acceptance criteria. Nothing in the process treated that exit as done: the task moved to awaiting verification, the gap was identified in review, and the work was redone. The routing rule and the run checks were tightened afterwards so the same gap would be caught earlier. This is the behavior the whole design exists for.

My contribution

Status and boundaries

Local, installable CLI · version 0.1 Single-stage executor Automatic scheduling not implemented

The CLI is a private preview package with a single-stage executor. The full flow is coordinated by a main session; automatic scheduling, persistent state, cross-process coordination, automatic review admission, and hard budget limits are future scope. I do not claim a measured cost reduction: the only figures I have are price scenarios under identical token and caching assumptions, not a controlled comparison.

Related public work

agent-harness-pack is the public, earlier half of this work: role templates for implementation and independent review, delivery rules, CI templates, and lessons learned. It is a separate repository from the CLI and is not the CLI's source code.