An AI workflow that
holds its shape
How the bigin-skills harness turns an AI coding assistant from
something that produces plausible code into something that produces
reviewable, specified, verified changes — and what you have to do to get that.
The problem this actually solves
Not "make the AI smarter" — make its output something you can review, and something the next session can pick up.
An unconstrained coding agent is good at producing code that looks right. That is the failure mode: the diff is plausible and large, so the ways it diverges from what you wanted surface only at review — when fixing them means a rewrite rather than a sentence.
Three specific things go wrong, and each has a mechanism in this plugin pointed at it:
- It starts before the shape is agreed. You get a large diff built on an assumption nobody stated. → the spec gate.
- It marks its own homework. "I implemented the feature and it works" is a self-report, not evidence. → an independent verifier that never sees that self-report.
- It forgets, expensively. Every session re-derives the same context, and re-derives it differently. → a knowledge bundle with an index-first read protocol.
What you get for it
Scoped to an approved plan, with an explicit "not in scope" line that makes creep visible.
Three execution tiers. Mechanical work doesn't pay for the top model, and you pick the ladder.
Deterministic scripts, not prose. Including one that stops the others being switched off.
What it costs you
The harness adds a decision point before non-trivial work: you read a spec and approve it. That is real friction, and it is the entire mechanism — an agent that never has to stop is an agent that never has to be right about scope. What it does not add is per-session overhead: the always-loaded footprint is a few thousand characters, and it is budget-checked at commit time so it stays that way.
Five concepts that explain the rest
Learn these and the individual skills stop needing separate explanation.
The spec gate
Before non-trivial work, the agent writes a short spec and waits for your
approval. Approved, it becomes PLAN.md — the artifact
everything downstream refers to. Bug fixes, copy changes, config tweaks and
changes under ~20 lines skip it entirely.
The two fields that carry the weight are Security considerations and Not in scope. The first is economics: a threat named at spec time is a sentence, the same threat found at review is a rewrite. The second is what makes "stop and ask" enforceable later — without it, scope creep has nothing to violate.
Independent verification
After implementation, a fresh verifier subagent audits the
diff against PLAN.md. It is passed the diff itself, never the
implementer's summary of what it did — that self-report is exactly what the
independence exists to route around. On a fail the original implementer is
resumed with the issue list; a new verifier checks the next round.
Capped at three rounds, then it stops and asks you. One exception, because a
resume is expensive and some failures are two words: where every issue names
the correct value, changes text rather than behaviour, and fits one bounded
hunk, the orchestrator applies the fix itself. A fresh verifier still audits
it, the round still counts against the three, and it is never done twice
running on the same issue.
Two axes, not one
Model choice answers "could it do this at all." Effort answers "did it check its work." They are scored separately, because a change can be mechanically trivial and still need heavy checking — a one-line edit to a contract, or a migration. Scoring them together means overpaying for the model or underpaying for the checking.
Where a fact lives
Six surfaces, one question each. Putting a fact in the wrong one is the most common way this convention degrades, because the wrong home has no mechanism to expire it.
| Surface | Answers | Lifetime |
|---|---|---|
| knowledge/ | What is the system, and why? | Outlives sprints; expires on behavior change |
| knowledge/implementation/ | How did this piece come to be, and what was rejected? | Append-only; never expires, never edited |
| docs/product/prd.md | What did we promise, and how is it checked? | Outlives every epic derived from it; amended, never renumbered |
| .claude/rules/ | How do we work here? | Outlives projects; changes by decision |
| graphify-out/ | Where is the code, what connects? | Regenerated; expires every commit |
| PLAN.md | What are we doing right now? | Archived verbatim at task end |
The always-loaded budget
Everything in CLAUDE.md and every unscoped rule file is read on
every turn of every session. So rules are scoped by path where
possible, and the total is checked by a script at commit time. This is why
the generated CLAUDE.md is short by design and why deep detail
lives in reference files that load only when needed.
Setting up a repo
One conversation, a handful of decisions, then one command each contributor runs once.
Install the plugin
/plugin marketplace add tammai/bigin-skills
/plugin install bigin-skills@bigin
npx skills add tammai/bigin-skills
/add-plugin
Then pick bigin-skills, or browse Customize → Plugins. The same skills/ and agents/ directories serve both hosts — the repo just carries a manifest for each. Cursor doesn't take the subagent ladder's per-tier model/effort pins, so model-router's fan-out stays Claude Code only; the skills, the workflow, and every gate work on both.
Run the harness setup
Say it in plain language — the skill triggers on intent, not a command name:
Set up the harness in this repo
You will be asked a small number of decisions. These are the ones worth thinking about:
| Decision | Recommendation |
|---|---|
| MODEL_ROUTING | Take the default opus-centric. Move later if verifier failures become routine. |
| KNOWLEDGE_BUNDLE | Yes for anything long-lived. It is the piece that compounds. |
| CI_PROVIDER | Yes if you have CI. The same checks run at commit and in CI, so they can't diverge. |
| GRAPH | Optional, and safe to defer. Most value on a large existing codebase. |
Activate the gates — once per clone, per contributor
ln -sf ../../scripts/pre-commit.sh .git/hooks/pre-commit && chmod +x scripts/pre-commit.sh
ln -sf ../../scripts/commit-msg.sh .git/hooks/commit-msg && chmod +x scripts/commit-msg.sh
simple-git-hooks or
husky already manages a hook, its own install step covers it and
setup leaves it alone. The exact pair for your repo is in the setup summary.
Read what it wrote
Open the generated CLAUDE.md. It is deliberately short, and it
is what every session in this repo sees before anything else. If something
in it is wrong for your team, fix it now — it is cheaper than any
correction you will make later.
Months later, when the repo has moved and the brief hasn't, run setup again and
choose verify. It installs nothing. It re-reads the
existing CLAUDE.md and checks each claim against the repo, running the
lint, typecheck and test commands the file names rather than trusting them — a
command that no longer exists is rewritten, a claim a manifest contradicts is
corrected, and anything it cannot verify is reported and left alone. It distinguishes
a command that cannot run from one that runs and fails: the second
means your suite is red, not that the brief is stale, and the row stays. A verify pass
may shrink the file or leave it the same size; one that grows it is a bug. Its sibling
patch is the other re-run mode — it applies only the
changes between the version your repo was scaffolded with and the current one.
The daily loop
Six steps. Two always need an answer from you, a third does when the work is big enough, and the rest run on their own.
/code-review, or
both — and only then is the plan distilled and archived.
Starting work
Say what you want. The workflow triggers on intent:
Add rate limiting to the public API endpoints
Fix the double-submit bug on the checkout form
You do not have to pick the right door. Every entry point triages through the same
ladder, so a request typed at the wrong one is handed to the right one rather than
spec'd over the gap. And if you would rather ask than guess,
/ask-bigin takes the request in plain language, names the one skill
that fits with the reason in a line, and hands off — it reads the same ladder for
build work, so it cannot send you somewhere the ladder wouldn't, and it never does
the work itself. It runs only when you type it: the skill sets
disable-model-invocation, so it can never wedge itself in front of a
request you already stated plainly.
What you do at each step
Scope and triage — read one sentence, before any code
What is changing, and why, in a single line. If it doesn't match what you meant, say so now: this is the cheapest correction available anywhere in the loop. The same step checks that sentence against the shared triage ladder, and only rung 1 continues — a request that turns out to be three plans stops here and goes to epic-workflow, one whose product shape isn't settled to discovery-workflow. Being invoked here is not evidence the work belongs here.
Spec gate — you read it and approve it
The one interruption that happens every time. Read the spec properly, and check Not in scope especially — it is the line that governs everything after it. If the request was underspecified you get up to three clarifying questions instead; answer them rather than waving them through.
PLAN.md — the working contract on disk
Nothing required from you. The approved spec plus a task table are written to disk, and the tasks are mirrored into the task list so you can watch progress without opening the file. PLAN.md stays the source of truth; the mirror is disposable.
Implement — the routed tier does the work
One possible confirmation. If scoring lands on the deep tier you are asked before it spawns — the most expensive tier and the biggest behavior swing, so it is the one pause worth having. Quick and standard spawn without asking.
Verify — a fresh agent, reading the diff only
Nothing required from you — this runs exactly as Independent verification describes. What to watch for is the third round: that is the cap, and it hands the decision back to you rather than looping again.
Review and Cleanup — one decision
Review first, and it is your call how. A human pass, /code-review, or both — it is offered, never run automatically. Say yes to the AI pass for anything touching auth, money, or data you can't reconstruct; a human pass is worth it wherever the consequence of being wrong lands on someone outside the team. If the change touches auth, sessions, secrets, PII, or untrusted input, /security-review is offered alongside it.
Cleanup only runs once review is resolved — clean, or explicitly declined. If the task established a decision, invariant, or constraint, a specific knowledge/ edit is proposed; then PLAN.md is archived — it stops being a working file, but it stays the only record of why the task took this shape, so it moves verbatim into knowledge/implementation/ (or .claude/memory/ in a repo with no bundle) rather than being deleted. "Nothing durable here" is the correct answer most of the time — routine work doesn't produce concepts.
When the requirement changes mid-task
You change your mind, or implementation reveals that a spec assumption was
wrong — and rows in PLAN.md are already done. Three things look
identical from inside the loop, and only the third one is this:
| What happened | What handles it |
|---|---|
| The code missed the plan | The verifier — a failing verdict, then a fix round |
| The task needs work outside its scope | Scope discipline — stop and ask, second task |
| The requirement moved | Course correction |
The plan freezes first. Its Status: line flips to
amending, and the spec gate — which only ever passes
approved — stops letting non-trivial edits through until you
re-approve. That is the existing gate doing its job, not a new one: nobody
builds against a spec that is mid-rewrite.
Then the change is classified out loud, because the three cases cost very different amounts. Additive means new rows and nothing already finished is invalidated. Invalidating means a finished row is now wrong, so it is flipped back with the reason recorded — a stale done row would make the verifier either fail honest work or pass work nobody asked for, because it audits the diff against the whole plan. A premise change means the goal itself moved, and the plan is replaced from the spec gate rather than patched; a plan patched past recognition reads as approved while describing something nobody agreed to.
Every amendment is logged to an ## Amendments section, which is the
only record that the plan's shape changed — the distill step reads it, and the
archive then carries it forward verbatim. Amendment rounds do not count
against the three-round fix-loop cap; they are a different failure mode. There
is a cap of two amendments per plan, after which the honest
conclusion is that the scope was wrong from the start, and it says so.
When the work is bigger than one task
Some requests are three or five plans, not one. With an approved PRD already on
disk, epic-workflow decomposes straight from its requirement index and
asks nothing — that file is the whole handoff from discovery, not a note to a human.
epic-workflow runs
above the loop and feeds it: it decides what the tasks are, in what
order, then hands them back one at a time.
task-workflow, and where the initiative is stated but
its product shape isn't settled it hands up to discovery-workflow.
Separately, above roughly eight units it calls the scope a roadmap and proposes the first epic-sized
slice instead. The return edge goes to the queue file, not to the session —
after /clear it re-reads the queue and takes the next eligible row.
Approving an epic approves the decomposition and nothing else. Every unit still faces the spec gate on its own merits. The queue file deliberately does not satisfy the guard — one epic-level approval standing in for five unwritten specs is precisely the drift the gate exists to stop.
One unit at a time, and it usually keeps going. Most of a
unit's weight never reaches your session — task-workflow's
implementer and verifier are both subagents — so what accumulates per unit is a
scope sentence, a spec, a plan, some verdicts and a review. Small, not nothing.
So the skill continues to the next unit by default and stops on a condition
instead of a schedule: when context is genuinely tight, when the finished unit
changed something a later unit inherits and you should see it before the next
spec is drafted against it, or when the next spec gate needs a decision the last
unit just changed. The queue file is the complete handoff package either way, so
a /clear whenever you like costs nothing. A genuinely independent
tail — rows with no Blocked by between them, touching disjoint
surfaces — can run as one worktree per unit instead; the skill names the rows and
stops rather than orchestrating that itself. The skill also refuses in both
directions of the ladder — below the bar back to the ordinary loop, and above it to
discovery-workflow when the product shape isn't settled. Separately,
above about eight units it says the scope is a roadmap, proposes the first
epic-sized slice, and names what it deferred.
Epic cleanup is where knowledge/ usually earns its keep. A single
plan rarely establishes anything durable; an epic that settled a contract or a
boundary did. Two things are written, and they must not restate each other: the
invariant the epic settled goes into a concept file, and
EPIC.md itself is archived verbatim as a record — goal, constraints,
the unit table with each unit's notes, and the amendment log. That notes column is
where each unit recorded what it changed for the units after it, which is precisely
what a distilled concept throws away.
When nobody can say what the thing is yet
Both layers above assume their subject arrives already stated — an
initiative, a task. discovery-workflow is the layer that
produces it. It runs when the ask is a vague one ("we want a client
portal"), a new product surface, or an initiative nobody can yet write a
single acceptance criterion for, and it stops at a PRD rather than going on
to decompose or implement.
docs/product/. The test between them is
not a feeling about size: can you write one testable acceptance criterion
for the request exactly as stated, inventing nothing? If you can, the
product question is already answered and this layer has no work to do.
Two homes, and nothing crosses between them. The brief and
the PRD are human-facing documents under docs/product/. The
decisions they force — an invariant, a boundary, a shape — are agent-facing
concept files under knowledge/architecture/. "A user only sees
their own workspace's invoices" is a promise somebody approves and somebody
later checks, so it is a requirement. "Every row belongs to one workspace and
every read filters by it" is a constraint on every future change that nobody
approves and everybody obeys, so it is a concept. Neither file restates the
other.
The PRD is a contract, which is why its shape is fixed.
Requirements are numbered blocks carrying a priority, the surfaces they
touch, what they depend on, and at least one acceptance criterion written so
it can be checked from outside the system. Numbers are assigned once and
never renumbered. Three things cite them: a full-spec PLAN.md's
Covers column, epic-workflow at its step 2, and
write-tests, which resolves an FR-3/AC-2 citation off disk
and quotes the criterion verbatim into an end-to-end spec. A tidy-up that closes a
gap silently re-points every
citation written against the old numbering, and citations are how the layers
below refer to a requirement at all.
It adds no gate either. The two approvals happen in the conversation, and nothing at commit time reads a brief or a PRD — a PRD gate would fire in every repo that installs the harness, and most of them will never have one. An approved artifact is never overwritten: re-invoking on a finished discovery reports what is on disk and goes straight to the handoff.
Knowledge that compounds
The part with the slowest payback and the highest ceiling.
Two kinds of concept, plus a record log
knowledge/ holds concept files answering what the
system is and why — decisions, contracts, boundaries, constraints
that outlive a sprint. Two populations of concept live there with different lifecycles:
- Your own system
- Written by the cleanup step and by end-of-sprint distillation. Tracks your code; expires when behavior changes. Verified by human review in the PR.
- A third-party library
- Written by
/knowledge-distillfrom the library's own repo at a pinned commit, then audited by a subagent against that clone. Expires when the dependency moves a minor version — enforced by a commit-time drift guard.
A third population sits on a different axis. knowledge/implementation/ is an
append-only record log — one file per finished task or epic, carrying the plan
verbatim with its task table and any amendments. Both populations above answer what the system
is and expire when behavior changes; a record answers how one piece came to be, and
never expires, because it was true when written. Records are never edited after they are
written, and they sit behind their own nested index, so they never load for routine work — you
go looking when you need to know why a past change took the shape it did.
The library case exists because models have a training cutoff and libraries don't. A pinned, audited bundle is what stops an agent confidently writing an API that was removed two versions ago.
knowledge/. "AuthService
calls TokenStore" is true until the next refactor and nothing will catch
it. Concept files point at sources of truth — the code, the
OpenAPI file, a rule — and never duplicate them.
The optional code graph
If you enable it, graphify parses your code into a graph of
symbols and their relationships, answering where questions
at zero token cost — what calls this, what breaks if I change it. It is
structure, never intent, and a source read always wins a disagreement.
The two are complements, and they fail differently. A stale concept file lies about intent, which only a human can catch — so it is defended by review and a validator. A stale graph lies about location, which a parser can re-derive in under a second — so it is defended by rebuilding.
When something blocks you
Each message names its own way out, and each one is real.
Gates are deterministic scripts running before a tool call or at commit time. Being blocked is the system working — the useful question is which gate, and what it wants.
| Message | What to do |
|---|---|
| PLAN.md missing or not approved | Approve a spec, or keep the change under 20 lines. Tests, docs and config are exempt. |
| PLAN.md is for branch 'X' | A leftover plan. Finish it, update its Branch: line, or delete it. |
| Not a Conventional Commit | Rewrite the subject: feat:, fix(api):, under 100 characters. |
| fix commit with no test | Stage the regression test, or add [no-test] with the reason. |
| --no-verify is blocked | Fix what is failing. This one has no override, deliberately. |
| Canary token detected | Stop. Not a false positive — see below. |
Why there is no bypass
--no-verify and force-pushes are blocked by a guard that runs
before the shell does. Every other gate is only as strong as the inability
to skip it, so that one is load-bearing. --force-with-lease is
allowed, because it refuses when the remote moved under you — which is the
actual damage --force does.
If a gate is wrong for your repo, change the gate — its allowlist, its threshold, the rule file — as a reviewed edit. That change shows up in a diff. A bypass doesn't.
The injection gate
Three stages defend against prompt injection arriving through fetched content. A scanner flags suspicious tool output; the next risky tool call asks for confirmation. Separately, a per-session random token is seeded with an instruction never to reproduce it — and any tool call containing it is denied outright. That third stage is not a heuristic: a session-unique UUID appearing in a tool call is proof the context is being exfiltrated. Treat the task as compromised and stop.
Best practices and recommendations
What actually determines whether this works for a team, in rough order of impact.
Do these
- Read the spec properly, once. Approving specs unread reproduces the exact problem the gate exists to prevent, with extra steps in between.
- Make everyone run the hooks-path command. Put it in your README's setup section. A contributor who skipped it has no gates and no way to notice.
- Start on the default ladder. Move to
frontieronly when verifier failures on standard-tier work become routine. One bad round is the loop working, not evidence of an under-powered tier. - Answer "nothing durable" honestly at cleanup. A bundle full of per-task narrative is worse than a small one, because the index stops being scannable.
- Fix the root cause upstream. When output is consistently wrong, the cause is usually an under-specified plan or a missing convention — not an under-powered model.
Avoid these
- Don't reach for the full spec because a task feels big. It is for an explicit request or a rubric score — and it auto-routes implementation to the most expensive tier, so casual use is expensive twice.
- Don't raise effort to fix quality. High is the ceiling by standing decision. The checking a higher setting would buy is already supplied by the verify loop; buying it twice yields slow, hedged output.
- Don't widen a gate's allowlist to unblock yourself. Especially not to include CI config — a workflow file executes with access to your secrets and defines which gates run at all.
- Don't let rules accumulate. Every addition to always-loaded context is paid on every turn of every session forever. End-of-sprint distillation enforces this: each addition must name what it replaces.
- Don't skip the branch line on a plan. It is what stops yesterday's plan silently governing today's task.
Rolling it out to a team
One repo first, for a sprint
Pick something real but not load-bearing. You want the gates to fire in anger before you have opinions about them.
Tune the rules, not the gates
Most friction is a convention that doesn't match how your team works. Fix it in .claude/rules/, where it is scoped and cheap.
Add the knowledge bundle when the first "why is it like this" appears
It compounds, but only if entries are earned. Starting it empty and letting the cleanup step feed it beats bulk-writing concepts up front.
Run an end-of-sprint distillation
This is the mechanism that keeps the harness from bloating over a year. It compresses rather than appends, and it stops for your approval before writing anything.
Appendix · Reference
Where to go when this handbook has stopped being enough.
Deeper guides
| Document | Covers |
|---|---|
| docs/USER_GUIDE.md | The task-oriented walkthrough, skill by skill |
| docs/SPEC-GATE.md | Both spec formats, and exactly what the guard measures |
| docs/ROUTING.md | Tier scoring, the three ladders, the verification bar |
| docs/GATES.md | Every guard, the injection gate, unblocking |
| docs/KNOWLEDGE.md | The bundle: what it holds, what keeps it honest |
| docs/GRAPHIFY.md | The optional code graph |
| README.md | Install, the skill and agent inventory, the model ladder |
Glossary
- Harness
- The generated files that govern AI work in a repo:
CLAUDE.md,.claude/rules/, and the commit-time gates. - Spec gate
- The requirement that non-trivial work has an approved
PLAN.mdbefore any edit lands. - Course correction
- The path taken when the requirement moves mid-task: the plan freezes at
Status: amendinguntil you re-approve the amended spec. Capped at two per plan. - PRD
- The approved requirements document at
docs/product/prd.md— numbered requirements, each with acceptance criteria. The contract the layers below cite by number. - Unit
- One row of an epic's queue — a single plan's worth of work, independently verifiable and shippable without any later unit existing.
- Tier
- Which subagent executes a task — quick, standard, or deep — chosen by capability scoring.
- Verification bar
- How carefully a change must be checked. Scored on its own axis rather than derived from the tier — it changes what the spawned agent is asked to demonstrate, not what the tier score was. One place it does move the outcome: a
quick-scored task with any bar trigger fired is spawned at the standard tier instead. - Ladder
- A profile mapping each tier to a model and an effort level. Three of them;
opus-centricis the default. - Concept file
- One Markdown file in
knowledge/describing a single invariant, contract, or domain idea. - Drift guard
- A commit-time check comparing a library bundle's pinned version against the dependency your project actually declares.
- Guard
- A deterministic Node script wired to a hook, blocking a tool call or a commit. Fails closed.
Getting unstuck
- A gate fired and you don't know why →
docs/GATES.md - Work routed to a tier you didn't expect → check which verification-bar trigger fired
- An agent wrote a stale library API → distil a pinned bundle for it
- Sessions keep re-deriving the same context → that is what
knowledge/is for