sous · Issue in, merged commit out
A software factory for coding agents, on your machine.
sous runs your software development flow with coding agents: implement, review, fix, merge, or any step you write. The flow is a deterministic graph in YAML, like a GitHub Actions workflow, so the agents do the work and the graph decides what happens next.
CLI first, with an optional live UI. Runs locally with the claude
or codex CLI you already have.
- deterministic flow graph
- runs locally, in a git worktree
- Claude Code or Codex, per step
- GitHub, beads or local issues
- open source, MIT
# in Claude Code: spec + sub-issues
> /sous-plan 57
# the rest runs without you
$ ./sous run task 57
$ ./sous status 57
# optional: watch it live
$ ./sous ui
The flow
Your development flow is a YAML file
A flow wires reusable steps, the starter's or your own, into a graph, the way a GitHub Actions workflow wires actions. Each entry names its step, the agent and model that runs it, its inputs and its routes: first match wins, a loop is a route back, and done and blocked are the only exits.
This is task, one of the example flows
sous init copies into your repository.
It is yours: change the steps, the routes and the models, or write flows that match your own cycle.
allowedLabels: [type:task, type:bug] # which issues this flow accepts
entry: init
steps:
- id: init # a module step: plain TypeScript, no agent
uses: init
routes:
- to: implement
- id: implement # an agent step: the flow names agent and model
uses: implement
agent: { kind: claude, model: claude-opus-5-5 }
with:
task: ${{ ctx.work.ref + " " + ctx.work.title + "\n\n" + ctx.work.body }}
routes:
- to: review
- id: review
uses: review
agent: { kind: claude, model: claude-opus-5-5 } # or a second opinion: { kind: codex, model: gpt-5.5 }
with:
diffRange: ${{ steps.implement.commits[0] + "^..HEAD" }}
criteria: ${{ ctx.work.body }}
routes: # ordered, first match wins
- to: sync
when: res.ok
- to: blocked # three rounds, then a human
when: steps.review.visits >= 3
- to: fix
else: true
- id: fix
uses: fix
agent: { kind: claude, model: claude-opus-5-5 }
resume: ${{ steps.fix.session ?? steps.implement.session }} # continue the agent's session
with:
input: ${{ steps.review.log }}
inputKind: review
…
routes:
- to: review # a loop is just a route back
- id: sync # bring main in; the agent runs only on a conflict
uses: sync-resolve
agent: { kind: claude, model: claude-opus-5-5 }
with:
baseBranch: main
routes:
- to: review-full # first time, or main moved
when: 'steps["review-full"] === undefined || prepared.output.moved || prepared.output.conflict'
- to: merge
else: true
- id: review-full # the whole branch against the spec
uses: review
agent: { kind: claude, model: claude-opus-5-5 }
with:
diffRange: main...HEAD
criteria: ${{ ctx.readFile(ctx.specDir + "/spec.md") }}
routes:
- to: sync
when: res.ok
- to: blocked
when: 'steps["review-full"].visits >= 4'
- to: fix-review-full
else: true
- id: fix-review-full
uses: fix
…
routes:
- to: review-full
- id: merge # squash onto main, label, close the issue
uses: merge
with:
baseBranch: main
mergedLabel: sous:merged
routes:
- to: done # done and blocked are the only exits
↓ THE SAME FLOW, MID-RUN, IN SOUS UI ON LOCALHOST:4680 · NOT A HOSTED SERVICE
./sous ui
draws every run as its flow, live, on localhost. It looks like GitHub Actions and runs on your machine. Here implement is done
and review is running, with each step's agent, model and duration. Scroll for the rest of the graph.
Why sous
An agent you can leave alone, because it never decides alone
Local-first
Like GitHub Actions, for agents, on your machine. Flows are YAML, steps are reusable units, routes make a graph.
Deterministic shell, agentic core
The agent works in a worktree and answers with one JSON envelope. The executor renders, evaluates and routes; it never asks the agent what happens next.
Quality gates
Evals judge the work, not the agent's word. A failed eval retries in the same session with the log; after the cap the run stops blocked instead of guessing.
Multi-model
A second opinion by design: each step names its agent and model, so Claude can implement and Codex can review.
Testability
Every step has declared inputs and an output schema, so you can run, measure and test it alone with fixed inputs. One big prompt or skill never gives you that. sous-factory/testing walks whole flows over fakes.
Observability
Every visit is recorded: inputs, agent log, eval verdict, commits. sous status prints the timeline, sous ui shows it live, and issue labels mirror the run's state.
Resumable
A blocked run keeps its worktree. Fix it by hand, run the same command, and it continues at that step's eval.
Customizable
sous init copies the steps and flows into your repository. From then on they're your code, with their own tests, not configuration.
Tracker-agnostic
GitHub out of the box, beads and local files built in, or a tracker module of your own.
The starter
Proven skills, two flows, yours to change
sous init copies a working factory into .sous/:
Matt Pocock's engineering skills
wired into two delivery flows. Start with it as it is, then change any step, prompt or route to match how your team builds software.
FLOWS
task- One issue. Implement → review ⇄ fix → sync main → full review → merge.
feature- A parent issue delivered sub-issue by sub-issue, then reviewed as a whole and merged.
Plus bench, a smoke-test flow that touches neither the issue nor the code.
SKILLS
implementandtddwrite the codecode-reviewjudges it, andtddfixes what it foundresolving-merge-conflictsbrings main into-spec,to-ticketsandgrillingplan the issue with you in/sous-plan
Design
Agentic core, deterministic shell
A step is a folder with a step.yaml: declared inputs and either a
prompt, where an agent does the work, or a TypeScript module. An agentic step can add a prepare, which may
skip the agent when there is nothing to do, and an eval, which judges the result.
A flow wires steps into a graph. Each entry names its step, its inputs, the agent and model, and ordered routes with conditions. done and blocked are the only exits; a loop is just a route back.
What this buys. Every decision that routes a run is a condition in YAML you can read, test and replay. Every judgement call lives in one prompt with one schema. A flaky agent cannot skip review, cannot merge, and cannot touch anything outside its worktree, because it was never given those tools.
Every step on the record
Each run writes a session record and an event log. sous ui
shows them live: the routes a step can take, each visit, the agent's log and the commits it made.
A run that cannot go on stops blocked and the worktree stays. Fix it by hand, run the same command, and the run continues at that step's eval.
Quick start
Running in your repository in minutes
Once per machine: Node 22, gh auth login, and claude
or codex on the PATH. Then, from the repository root:
SET UP
npm install -D sous-factory
npx sous init # copies the starter into .sous/
cd .sous
./sous validate
./sous run bench <issue> # smoke test, touches nothing
DAILY USE
/sous-plan <issue> # in Claude Code
./sous run feature <issue> # or task
./sous status <issue> # facts + timeline
./sous graph task # the flow as a diagram
./sous ui # live, on localhost
sous init leaves a .sous/ folder with the flows, steps,
skills, prompts and tests of your factory, all yours to edit. Prefer a guided setup? The /sous-setup
skill installs, inits and adapts the files with you in Claude Code.
A person plans. Agents build in worktrees. A deterministic executor runs every check. Only what passes is merged.