sous · Issue in, merged commit out

A software factory for coding agents, on your machine.

sous runs your software development flow with coding agents: implement, review, fix, merge, or any step you write. The flow is a deterministic graph in YAML, like a GitHub Actions workflow, so the agents do the work and the graph decides what happens next.

CLI first, with an optional live UI. Runs locally with the claude or codex CLI you already have.

  • deterministic flow graph
  • runs locally, in a git worktree
  • Claude Code or Codex, per step
  • GitHub, beads or local issues
  • open source, MIT
npm install -D sous-factory
See a flow ↓

Node 22 · GitHub · npm

acme-web
# in Claude Code: spec + sub-issues
> /sous-plan 57

# the rest runs without you
$ ./sous run task 57
$ ./sous status 57

# optional: watch it live
$ ./sous ui

The flow

Your development flow is a YAML file

A flow wires reusable steps, the starter's or your own, into a graph, the way a GitHub Actions workflow wires actions. Each entry names its step, the agent and model that runs it, its inputs and its routes: first match wins, a loop is a route back, and done and blocked are the only exits.

This is task, one of the example flows sous init copies into your repository. It is yours: change the steps, the routes and the models, or write flows that match your own cycle.

.sous/flows/task.yamlannotated · scroll ↓
allowedLabels: [type:task, type:bug]       # which issues this flow accepts
entry: init
steps:
  - id: init                               # a module step: plain TypeScript, no agent
    uses: init
    routes:
      - to: implement

  - id: implement                          # an agent step: the flow names agent and model
    uses: implement
    agent: { kind: claude, model: claude-opus-5-5 }
    with:
      task: ${{ ctx.work.ref + " " + ctx.work.title + "\n\n" + ctx.work.body }}
    routes:
      - to: review

  - id: review
    uses: review
    agent: { kind: claude, model: claude-opus-5-5 }   # or a second opinion: { kind: codex, model: gpt-5.5 }
    with:
      diffRange: ${{ steps.implement.commits[0] + "^..HEAD" }}
      criteria:  ${{ ctx.work.body }}
    routes:                                # ordered, first match wins
      - to: sync
        when: res.ok
      - to: blocked                        # three rounds, then a human
        when: steps.review.visits >= 3
      - to: fix
        else: true

  - id: fix
    uses: fix
    agent: { kind: claude, model: claude-opus-5-5 }
    resume: ${{ steps.fix.session ?? steps.implement.session }}   # continue the agent's session
    with:
      input:     ${{ steps.review.log }}
      inputKind: review
      …
    routes:
      - to: review                         # a loop is just a route back

  - id: sync                               # bring main in; the agent runs only on a conflict
    uses: sync-resolve
    agent: { kind: claude, model: claude-opus-5-5 }
    with:
      baseBranch: main
    routes:
      - to: review-full                    # first time, or main moved
        when: 'steps["review-full"] === undefined || prepared.output.moved || prepared.output.conflict'
      - to: merge
        else: true

  - id: review-full                        # the whole branch against the spec
    uses: review
    agent: { kind: claude, model: claude-opus-5-5 }
    with:
      diffRange: main...HEAD
      criteria:  ${{ ctx.readFile(ctx.specDir + "/spec.md") }}
    routes:
      - to: sync
        when: res.ok
      - to: blocked
        when: 'steps["review-full"].visits >= 4'
      - to: fix-review-full
        else: true

  - id: fix-review-full
    uses: fix
    …
    routes:
      - to: review-full

  - id: merge                              # squash onto main, label, close the issue
    uses: merge
    with:
      baseBranch: main
      mergedLabel: sous:merged
    routes:
      - to: done                           # done and blocked are the only exits

↓ THE SAME FLOW, MID-RUN, IN SOUS UI ON LOCALHOST:4680 · NOT A HOSTED SERVICE

sous ui flow graph for task 57: init and implement done, review running; fix, sync, review-full, fix-review-full and merge still queued
Optional UI. sous is a CLI first; ./sous ui draws every run as its flow, live, on localhost. It looks like GitHub Actions and runs on your machine. Here implement is done and review is running, with each step's agent, model and duration. Scroll for the rest of the graph.

Why sous

An agent you can leave alone, because it never decides alone

Local-first

Like GitHub Actions, for agents, on your machine. Flows are YAML, steps are reusable units, routes make a graph.

Deterministic shell, agentic core

The agent works in a worktree and answers with one JSON envelope. The executor renders, evaluates and routes; it never asks the agent what happens next.

Quality gates

Evals judge the work, not the agent's word. A failed eval retries in the same session with the log; after the cap the run stops blocked instead of guessing.

Multi-model

A second opinion by design: each step names its agent and model, so Claude can implement and Codex can review.

Testability

Every step has declared inputs and an output schema, so you can run, measure and test it alone with fixed inputs. One big prompt or skill never gives you that. sous-factory/testing walks whole flows over fakes.

Observability

Every visit is recorded: inputs, agent log, eval verdict, commits. sous status prints the timeline, sous ui shows it live, and issue labels mirror the run's state.

Resumable

A blocked run keeps its worktree. Fix it by hand, run the same command, and it continues at that step's eval.

Customizable

sous init copies the steps and flows into your repository. From then on they're your code, with their own tests, not configuration.

Tracker-agnostic

GitHub out of the box, beads and local files built in, or a tracker module of your own.

The starter

Proven skills, two flows, yours to change

sous init copies a working factory into .sous/: Matt Pocock's engineering skills wired into two delivery flows. Start with it as it is, then change any step, prompt or route to match how your team builds software.

FLOWS

task
One issue. Implement → review ⇄ fix → sync main → full review → merge.
feature
A parent issue delivered sub-issue by sub-issue, then reviewed as a whole and merged.

Plus bench, a smoke-test flow that touches neither the issue nor the code.

SKILLS

  • implement and tdd write the code
  • code-review judges it, and tdd fixes what it found
  • resolving-merge-conflicts brings main in
  • to-spec, to-tickets and grilling plan the issue with you in /sous-plan

Design

Agentic core, deterministic shell

A step is a folder with a step.yaml: declared inputs and either a prompt, where an agent does the work, or a TypeScript module. An agentic step can add a prepare, which may skip the agent when there is nothing to do, and an eval, which judges the result.

A flow wires steps into a graph. Each entry names its step, its inputs, the agent and model, and ordered routes with conditions. done and blocked are the only exits; a loop is just a route back.

What this buys. Every decision that routes a run is a condition in YAML you can read, test and replay. Every judgement call lives in one prompt with one schema. A flaky agent cannot skip review, cannot merge, and cannot touch anything outside its worktree, because it was never given those tools.

EXECUTOR · DETERMINISTIC SHELL TypeScript, no model, every branch a tested condition prepare · render run the agent eval → { ok, log } route ok retry ≤ max next entry blocked cap WORKTREE one git worktree on the run branch, nothing else agent claude · codex CLI edits, runs, commits prompt + declared inputs { summary, commits, output, skills } answers with one JSON envelope no routing, no merge, no say in the eval
One flow entry. Routes are ordered conditions, first match wins. The executor owns the loop; the agent sees only a rendered prompt and its worktree, and hands back a fixed envelope. Judging happens outside the agent, so a confident agent cannot approve its own work.

Every step on the record

Each run writes a session record and an event log. sous ui shows them live: the routes a step can take, each visit, the agent's log and the commits it made.

A run that cannot go on stops blocked and the worktree stays. Fix it by hand, run the same command, and the run continues at that step's eval.

sous ui step view for implement: routes with first match wins, the current visit, the running agent and its output
A step's view: its routes, the current visit and the agent log.

Quick start

Running in your repository in minutes

Once per machine: Node 22, gh auth login, and claude or codex on the PATH. Then, from the repository root:

SET UP

npm install -D sous-factory
npx sous init        # copies the starter into .sous/

cd .sous
./sous validate
./sous run bench <issue>   # smoke test, touches nothing

DAILY USE

/sous-plan <issue>          # in Claude Code
./sous run feature <issue>  # or task
./sous status <issue>       # facts + timeline
./sous graph task           # the flow as a diagram
./sous ui                   # live, on localhost

sous init leaves a .sous/ folder with the flows, steps, skills, prompts and tests of your factory, all yours to edit. Prefer a guided setup? The /sous-setup skill installs, inits and adapts the files with you in Claude Code.

A person plans. Agents build in worktrees. A deterministic executor runs every check. Only what passes is merged.

npm install -D sous-factory