BigIn · Engineering Handbook

An AI workflow that
holds its shape

How the bigin-skills harness turns an AI coding assistant from something that produces plausible code into something that produces reviewable, specified, verified changes — and what you have to do to get that.

Start here ~30 minutes Any repo Nuxt · Next · Go · Node · Flutter
BigIn · 2026
Part 1 · Orientation

The problem this actually solves

Not "make the AI smarter" — make its output something you can review, and something the next session can pick up.

An unconstrained coding agent is good at producing code that looks right. That is the failure mode: the diff is plausible and large, so the ways it diverges from what you wanted surface only at review — when fixing them means a rewrite rather than a sentence.

Three specific things go wrong, and each has a mechanism in this plugin pointed at it:

  • It starts before the shape is agreed. You get a large diff built on an assumption nobody stated. → the spec gate.
  • It marks its own homework. "I implemented the feature and it works" is a self-report, not evidence. → an independent verifier that never sees that self-report.
  • It forgets, expensively. Every session re-derives the same context, and re-derives it differently. → a knowledge bundle with an index-first read protocol.

What you get for it

Reviewable diffs

Scoped to an approved plan, with an explicit "not in scope" line that makes creep visible.

Cost you control

Three execution tiers. Mechanical work doesn't pay for the top model, and you pick the ladder.

Gates that hold

Deterministic scripts, not prose. Including one that stops the others being switched off.

Key idea Every rule here is enforced by something that isn't the model. A rule an agent can talk itself out of is a suggestion, and suggestions decay.

What it costs you

The harness adds a decision point before non-trivial work: you read a spec and approve it. That is real friction, and it is the entire mechanism — an agent that never has to stop is an agent that never has to be right about scope. What it does not add is per-session overhead: the always-loaded footprint is a few thousand characters, and it is budget-checked at commit time so it stays that way.

Five concepts that explain the rest

Learn these and the individual skills stop needing separate explanation.

The spec gate

Before non-trivial work, the agent writes a short spec and waits for your approval. Approved, it becomes PLAN.md — the artifact everything downstream refers to. Bug fixes, copy changes, config tweaks and changes under ~20 lines skip it entirely.

The two fields that carry the weight are Security considerations and Not in scope. The first is economics: a threat named at spec time is a sentence, the same threat found at review is a rewrite. The second is what makes "stop and ask" enforceable later — without it, scope creep has nothing to violate.

Independent verification

After implementation, a fresh verifier subagent audits the diff against PLAN.md. It is passed the diff itself, never the implementer's summary of what it did — that self-report is exactly what the independence exists to route around. On a fail the original implementer is resumed with the issue list; a new verifier checks the next round. Capped at three rounds, then it stops and asks you. One exception, because a resume is expensive and some failures are two words: where every issue names the correct value, changes text rather than behaviour, and fits one bounded hunk, the orchestrator applies the fix itself. A fresh verifier still audits it, the round still counts against the three, and it is never done twice running on the same issue.

Two axes, not one

Model choice answers "could it do this at all." Effort answers "did it check its work." They are scored separately, because a change can be mechanically trivial and still need heavy checking — a one-line edit to a contract, or a migration. Scoring them together means overpaying for the model or underpaying for the checking.

Where a fact lives

Six surfaces, one question each. Putting a fact in the wrong one is the most common way this convention degrades, because the wrong home has no mechanism to expire it.

SurfaceAnswersLifetime
knowledge/What is the system, and why?Outlives sprints; expires on behavior change
knowledge/implementation/How did this piece come to be, and what was rejected?Append-only; never expires, never edited
docs/product/prd.mdWhat did we promise, and how is it checked?Outlives every epic derived from it; amended, never renumbered
.claude/rules/How do we work here?Outlives projects; changes by decision
graphify-out/Where is the code, what connects?Regenerated; expires every commit
PLAN.mdWhat are we doing right now?Archived verbatim at task end

The always-loaded budget

Everything in CLAUDE.md and every unscoped rule file is read on every turn of every session. So rules are scoped by path where possible, and the total is checked by a script at commit time. This is why the generated CLAUDE.md is short by design and why deep detail lives in reference files that load only when needed.

Note The same principle drives the knowledge bundle's index-first read protocol: the bundle can grow for years without growing what a session loads, because the default read is the one-line index, not the contents.
Part 2 · Getting started

Setting up a repo

One conversation, a handful of decisions, then one command each contributor runs once.

Install the plugin

claude code
/plugin marketplace add tammai/bigin-skills
/plugin install bigin-skills@bigin
terminal
npx skills add tammai/bigin-skills
cursor
/add-plugin

Then pick bigin-skills, or browse Customize → Plugins. The same skills/ and agents/ directories serve both hosts — the repo just carries a manifest for each. Cursor doesn't take the subagent ladder's per-tier model/effort pins, so model-router's fan-out stays Claude Code only; the skills, the workflow, and every gate work on both.

Run the harness setup

Say it in plain language — the skill triggers on intent, not a command name:

prompt
Set up the harness in this repo

You will be asked a small number of decisions. These are the ones worth thinking about:

DecisionRecommendation
MODEL_ROUTINGTake the default opus-centric. Move later if verifier failures become routine.
KNOWLEDGE_BUNDLEYes for anything long-lived. It is the piece that compounds.
CI_PROVIDERYes if you have CI. The same checks run at commit and in CI, so they can't diverge.
GRAPHOptional, and safe to defer. Most value on a large existing codebase.

Activate the gates — once per clone, per contributor

terminal
ln -sf ../../scripts/pre-commit.sh .git/hooks/pre-commit && chmod +x scripts/pre-commit.sh
ln -sf ../../scripts/commit-msg.sh .git/hooks/commit-msg && chmod +x scripts/commit-msg.sh
Watch out Git hooks are not cloned. Skip this and every commit-time gate is silently inactive for that person — with no error and no visible difference. It is the single most common setup miss. Skip either line whose script your repo doesn't have: where simple-git-hooks or husky already manages a hook, its own install step covers it and setup leaves it alone. The exact pair for your repo is in the setup summary.

Read what it wrote

Open the generated CLAUDE.md. It is deliberately short, and it is what every session in this repo sees before anything else. If something in it is wrong for your team, fix it now — it is cheaper than any correction you will make later.

Months later, when the repo has moved and the brief hasn't, run setup again and choose verify. It installs nothing. It re-reads the existing CLAUDE.md and checks each claim against the repo, running the lint, typecheck and test commands the file names rather than trusting them — a command that no longer exists is rewritten, a claim a manifest contradicts is corrected, and anything it cannot verify is reported and left alone. It distinguishes a command that cannot run from one that runs and fails: the second means your suite is red, not that the brief is stale, and the row stays. A verify pass may shrink the file or leave it the same size; one that grows it is a bug. Its sibling patch is the other re-run mode — it applies only the changes between the version your repo was scaffolded with and the current one.

The daily loop

Six steps. Two always need an answer from you, a third does when the work is big enough, and the rest run on their own.

Scope one sentence, before any code Spec gate you read it and approve it PLAN.md the working contract on disk Implement the routed tier does the work Verify fresh agent, reads the diff only Review and Cleanup by a human or /code-review then distill and archive the plan
The dashed return is the fix loop: a failing verdict resumes the same implementer — or, for a trivially textual fix, is applied in place — and a brand-new verifier checks the next round either way. Capped at three. Review is your call — a human pass, /code-review, or both — and only then is the plan distilled and archived.

Starting work

Say what you want. The workflow triggers on intent:

prompt
Add rate limiting to the public API endpoints
Fix the double-submit bug on the checkout form

You do not have to pick the right door. Every entry point triages through the same ladder, so a request typed at the wrong one is handed to the right one rather than spec'd over the gap. And if you would rather ask than guess, /ask-bigin takes the request in plain language, names the one skill that fits with the reason in a line, and hands off — it reads the same ladder for build work, so it cannot send you somewhere the ladder wouldn't, and it never does the work itself. It runs only when you type it: the skill sets disable-model-invocation, so it can never wedge itself in front of a request you already stated plainly.

What you do at each step

Scope and triage — read one sentence, before any code

What is changing, and why, in a single line. If it doesn't match what you meant, say so now: this is the cheapest correction available anywhere in the loop. The same step checks that sentence against the shared triage ladder, and only rung 1 continues — a request that turns out to be three plans stops here and goes to epic-workflow, one whose product shape isn't settled to discovery-workflow. Being invoked here is not evidence the work belongs here.

Spec gate — you read it and approve it

The one interruption that happens every time. Read the spec properly, and check Not in scope especially — it is the line that governs everything after it. If the request was underspecified you get up to three clarifying questions instead; answer them rather than waving them through.

PLAN.md — the working contract on disk

Nothing required from you. The approved spec plus a task table are written to disk, and the tasks are mirrored into the task list so you can watch progress without opening the file. PLAN.md stays the source of truth; the mirror is disposable.

Implement — the routed tier does the work

One possible confirmation. If scoring lands on the deep tier you are asked before it spawns — the most expensive tier and the biggest behavior swing, so it is the one pause worth having. Quick and standard spawn without asking.

Verify — a fresh agent, reading the diff only

Nothing required from you — this runs exactly as Independent verification describes. What to watch for is the third round: that is the cap, and it hands the decision back to you rather than looping again.

Review and Cleanup — one decision

Review first, and it is your call how. A human pass, /code-review, or both — it is offered, never run automatically. Say yes to the AI pass for anything touching auth, money, or data you can't reconstruct; a human pass is worth it wherever the consequence of being wrong lands on someone outside the team. If the change touches auth, sessions, secrets, PII, or untrusted input, /security-review is offered alongside it.

Cleanup only runs once review is resolved — clean, or explicitly declined. If the task established a decision, invariant, or constraint, a specific knowledge/ edit is proposed; then PLAN.md is archived — it stops being a working file, but it stays the only record of why the task took this shape, so it moves verbatim into knowledge/implementation/ (or .claude/memory/ in a repo with no bundle) rather than being deleted. "Nothing durable here" is the correct answer most of the time — routine work doesn't produce concepts.

When the requirement changes mid-task

You change your mind, or implementation reveals that a spec assumption was wrong — and rows in PLAN.md are already done. Three things look identical from inside the loop, and only the third one is this:

What happenedWhat handles it
The code missed the planThe verifier — a failing verdict, then a fix round
The task needs work outside its scopeScope discipline — stop and ask, second task
The requirement movedCourse correction

The plan freezes first. Its Status: line flips to amending, and the spec gate — which only ever passes approved — stops letting non-trivial edits through until you re-approve. That is the existing gate doing its job, not a new one: nobody builds against a spec that is mid-rewrite.

Then the change is classified out loud, because the three cases cost very different amounts. Additive means new rows and nothing already finished is invalidated. Invalidating means a finished row is now wrong, so it is flipped back with the reason recorded — a stale done row would make the verifier either fail honest work or pass work nobody asked for, because it audits the diff against the whole plan. A premise change means the goal itself moved, and the plan is replaced from the spec gate rather than patched; a plan patched past recognition reads as approved while describing something nobody agreed to.

Every amendment is logged to an ## Amendments section, which is the only record that the plan's shape changed — the distill step reads it, and the archive then carries it forward verbatim. Amendment rounds do not count against the three-round fix-loop cap; they are a different failure mode. There is a cap of two amendments per plan, after which the honest conclusion is that the scope was wrong from the start, and it says so.

When the work is bigger than one task

Some requests are three or five plans, not one. With an approved PRD already on disk, epic-workflow decomposes straight from its requirement index and asks nothing — that file is the whole handoff from discovery, not a note to a human. epic-workflow runs above the loop and feeds it: it decides what the tasks are, in what order, then hands them back one at a time.

An initiative Clears the bar? 3+ units · >1 PR · 2+ surfaces no Down to task-workflow one plan, one PR yes Product shape settled? can you write one testable criterion? no Up to discovery-workflow it ends by handing back a PRD yes Up to 3 questions — or zero zero when an approved PRD is on disk; else about the decomposition, not the code Decompose into ordered units each one plan's worth, each shippable alone Above ~8 units a roadmap — first slice only You approve the decomposition the epic's only gate — nothing on disk before it The queue is written .claude/memory/EPIC.md One unit at a time the six-step loop above runs inside row done → next unit /clear only when it stops every row done Distill, then archive the queue the invariant to knowledge/, the queue verbatim
The skill refuses in both directions of the ladder: below the bar it hands the request back to task-workflow, and where the initiative is stated but its product shape isn't settled it hands up to discovery-workflow. Separately, above roughly eight units it calls the scope a roadmap and proposes the first epic-sized slice instead. The return edge goes to the queue file, not to the session — after /clear it re-reads the queue and takes the next eligible row.

Approving an epic approves the decomposition and nothing else. Every unit still faces the spec gate on its own merits. The queue file deliberately does not satisfy the guard — one epic-level approval standing in for five unwritten specs is precisely the drift the gate exists to stop.

One unit at a time, and it usually keeps going. Most of a unit's weight never reaches your session — task-workflow's implementer and verifier are both subagents — so what accumulates per unit is a scope sentence, a spec, a plan, some verdicts and a review. Small, not nothing. So the skill continues to the next unit by default and stops on a condition instead of a schedule: when context is genuinely tight, when the finished unit changed something a later unit inherits and you should see it before the next spec is drafted against it, or when the next spec gate needs a decision the last unit just changed. The queue file is the complete handoff package either way, so a /clear whenever you like costs nothing. A genuinely independent tail — rows with no Blocked by between them, touching disjoint surfaces — can run as one worktree per unit instead; the skill names the rows and stops rather than orchestrating that itself. The skill also refuses in both directions of the ladder — below the bar back to the ordinary loop, and above it to discovery-workflow when the product shape isn't settled. Separately, above about eight units it says the scope is a roadmap, proposes the first epic-sized slice, and names what it deferred.

Epic cleanup is where knowledge/ usually earns its keep. A single plan rarely establishes anything durable; an epic that settled a contract or a boundary did. Two things are written, and they must not restate each other: the invariant the epic settled goes into a concept file, and EPIC.md itself is archived verbatim as a record — goal, constraints, the unit table with each unit's notes, and the amendment log. That notes column is where each unit recorded what it changed for the units after it, which is precisely what a distilled concept throws away.

When nobody can say what the thing is yet

Both layers above assume their subject arrives already stated — an initiative, a task. discovery-workflow is the layer that produces it. It runs when the ask is a vague one ("we want a client portal"), a new product surface, or an initiative nobody can yet write a single acceptance criterion for, and it stops at a PRD rather than going on to decompose or implement.

A vague product ask Which rung? one criterion, inventing nothing? A change or a bug task-workflow · nothing written A settled initiative epic-workflow · nothing written neither Read the repo before asking every command run before it is written down Elicit only what it cannot answer 3 rounds, 12 questions, then it writes anyway You approve the brief nothing on disk before this docs/product/brief.md You approve the PRD roles, never people — the privacy checkpoint docs/product/prd.md FR-n requirements, acceptance criteria each The decisions the product forced knowledge/architecture/ skipped, loudly, with no bundle Hand the PRD path to epic-workflow then stop — the file is the whole handoff package
Both early exits are the common case, and both write nothing at all — no stub brief, no empty docs/product/. The test between them is not a feeling about size: can you write one testable acceptance criterion for the request exactly as stated, inventing nothing? If you can, the product question is already answered and this layer has no work to do.

Two homes, and nothing crosses between them. The brief and the PRD are human-facing documents under docs/product/. The decisions they force — an invariant, a boundary, a shape — are agent-facing concept files under knowledge/architecture/. "A user only sees their own workspace's invoices" is a promise somebody approves and somebody later checks, so it is a requirement. "Every row belongs to one workspace and every read filters by it" is a constraint on every future change that nobody approves and everybody obeys, so it is a concept. Neither file restates the other.

The PRD is a contract, which is why its shape is fixed. Requirements are numbered blocks carrying a priority, the surfaces they touch, what they depend on, and at least one acceptance criterion written so it can be checked from outside the system. Numbers are assigned once and never renumbered. Three things cite them: a full-spec PLAN.md's Covers column, epic-workflow at its step 2, and write-tests, which resolves an FR-3/AC-2 citation off disk and quotes the criterion verbatim into an end-to-end spec. A tidy-up that closes a gap silently re-points every citation written against the old numbering, and citations are how the layers below refer to a requirement at all.

It adds no gate either. The two approvals happen in the conversation, and nothing at commit time reads a brief or a PRD — a PRD gate would fire in every repo that installs the harness, and most of them will never have one. An approved artifact is never overwritten: re-invoking on a finished discovery reports what is on disk and goes straight to the handoff.

Note Skipping this layer entirely is fine and usually right. A repo with no PRD behaves exactly as it always has — the layer exists for the case where the first honest answer to "what are we building?" is "we are not sure yet," which is a bad thing to discover three units into an epic.
Tip Scope discipline is a hard rule, not a preference: if implementation reveals work outside the stated scope, the agent stops and asks. A second task is better than a sprawling first one — and if that second task turns out to be one of several, the epic layer above takes over rather than the plan growing.
Part 3 · Living with it

Knowledge that compounds

The part with the slowest payback and the highest ceiling.

Two kinds of concept, plus a record log

knowledge/ holds concept files answering what the system is and why — decisions, contracts, boundaries, constraints that outlive a sprint. Two populations of concept live there with different lifecycles:

Your own system
Written by the cleanup step and by end-of-sprint distillation. Tracks your code; expires when behavior changes. Verified by human review in the PR.
A third-party library
Written by /knowledge-distill from the library's own repo at a pinned commit, then audited by a subagent against that clone. Expires when the dependency moves a minor version — enforced by a commit-time drift guard.

A third population sits on a different axis. knowledge/implementation/ is an append-only record log — one file per finished task or epic, carrying the plan verbatim with its task table and any amendments. Both populations above answer what the system is and expire when behavior changes; a record answers how one piece came to be, and never expires, because it was true when written. Records are never edited after they are written, and they sit behind their own nested index, so they never load for routine work — you go looking when you need to know why a past change took the shape it did.

The library case exists because models have a training cutoff and libraries don't. A pinned, audited bundle is what stops an agent confidently writing an API that was removed two versions ago.

Watch out Never restate structural facts in knowledge/. "AuthService calls TokenStore" is true until the next refactor and nothing will catch it. Concept files point at sources of truth — the code, the OpenAPI file, a rule — and never duplicate them.

The optional code graph

If you enable it, graphify parses your code into a graph of symbols and their relationships, answering where questions at zero token cost — what calls this, what breaks if I change it. It is structure, never intent, and a source read always wins a disagreement.

The two are complements, and they fail differently. A stale concept file lies about intent, which only a human can catch — so it is defended by review and a validator. A stale graph lies about location, which a parser can re-derive in under a second — so it is defended by rebuilding.

When something blocks you

Each message names its own way out, and each one is real.

Gates are deterministic scripts running before a tool call or at commit time. Being blocked is the system working — the useful question is which gate, and what it wants.

MessageWhat to do
PLAN.md missing or not approvedApprove a spec, or keep the change under 20 lines. Tests, docs and config are exempt.
PLAN.md is for branch 'X'A leftover plan. Finish it, update its Branch: line, or delete it.
Not a Conventional CommitRewrite the subject: feat:, fix(api):, under 100 characters.
fix commit with no testStage the regression test, or add [no-test] with the reason.
--no-verify is blockedFix what is failing. This one has no override, deliberately.
Canary token detectedStop. Not a false positive — see below.

Why there is no bypass

--no-verify and force-pushes are blocked by a guard that runs before the shell does. Every other gate is only as strong as the inability to skip it, so that one is load-bearing. --force-with-lease is allowed, because it refuses when the remote moved under you — which is the actual damage --force does.

If a gate is wrong for your repo, change the gate — its allowlist, its threshold, the rule file — as a reviewed edit. That change shows up in a diff. A bypass doesn't.

The injection gate

Three stages defend against prompt injection arriving through fetched content. A scanner flags suspicious tool output; the next risky tool call asks for confirmation. Separately, a per-session random token is seeded with an instruction never to reproduce it — and any tool call containing it is denied outright. That third stage is not a heuristic: a session-unique UUID appearing in a tool call is proof the context is being exfiltrated. Treat the task as compromised and stop.

Best practices and recommendations

What actually determines whether this works for a team, in rough order of impact.

Do these

  • Read the spec properly, once. Approving specs unread reproduces the exact problem the gate exists to prevent, with extra steps in between.
  • Make everyone run the hooks-path command. Put it in your README's setup section. A contributor who skipped it has no gates and no way to notice.
  • Start on the default ladder. Move to frontier only when verifier failures on standard-tier work become routine. One bad round is the loop working, not evidence of an under-powered tier.
  • Answer "nothing durable" honestly at cleanup. A bundle full of per-task narrative is worse than a small one, because the index stops being scannable.
  • Fix the root cause upstream. When output is consistently wrong, the cause is usually an under-specified plan or a missing convention — not an under-powered model.

Avoid these

  • Don't reach for the full spec because a task feels big. It is for an explicit request or a rubric score — and it auto-routes implementation to the most expensive tier, so casual use is expensive twice.
  • Don't raise effort to fix quality. High is the ceiling by standing decision. The checking a higher setting would buy is already supplied by the verify loop; buying it twice yields slow, hedged output.
  • Don't widen a gate's allowlist to unblock yourself. Especially not to include CI config — a workflow file executes with access to your secrets and defines which gates run at all.
  • Don't let rules accumulate. Every addition to always-loaded context is paid on every turn of every session forever. End-of-sprint distillation enforces this: each addition must name what it replaces.
  • Don't skip the branch line on a plan. It is what stops yesterday's plan silently governing today's task.

Rolling it out to a team

One repo first, for a sprint

Pick something real but not load-bearing. You want the gates to fire in anger before you have opinions about them.

Tune the rules, not the gates

Most friction is a convention that doesn't match how your team works. Fix it in .claude/rules/, where it is scoped and cheap.

Add the knowledge bundle when the first "why is it like this" appears

It compounds, but only if entries are earned. Starting it empty and letting the cleanup step feed it beats bulk-writing concepts up front.

Run an end-of-sprint distillation

This is the mechanism that keeps the harness from bloating over a year. It compresses rather than appends, and it stops for your approval before writing anything.

Key idea The harness is not a productivity multiplier on its own — it is a consistency mechanism. Its value shows up in the changes that didn't need a second pass, and in the sessions that started already knowing what the last one learned.

Appendix · Reference

Where to go when this handbook has stopped being enough.

Deeper guides

DocumentCovers
docs/USER_GUIDE.mdThe task-oriented walkthrough, skill by skill
docs/SPEC-GATE.mdBoth spec formats, and exactly what the guard measures
docs/ROUTING.mdTier scoring, the three ladders, the verification bar
docs/GATES.mdEvery guard, the injection gate, unblocking
docs/KNOWLEDGE.mdThe bundle: what it holds, what keeps it honest
docs/GRAPHIFY.mdThe optional code graph
README.mdInstall, the skill and agent inventory, the model ladder

Glossary

Harness
The generated files that govern AI work in a repo: CLAUDE.md, .claude/rules/, and the commit-time gates.
Spec gate
The requirement that non-trivial work has an approved PLAN.md before any edit lands.
Course correction
The path taken when the requirement moves mid-task: the plan freezes at Status: amending until you re-approve the amended spec. Capped at two per plan.
PRD
The approved requirements document at docs/product/prd.md — numbered requirements, each with acceptance criteria. The contract the layers below cite by number.
Unit
One row of an epic's queue — a single plan's worth of work, independently verifiable and shippable without any later unit existing.
Tier
Which subagent executes a task — quick, standard, or deep — chosen by capability scoring.
Verification bar
How carefully a change must be checked. Scored on its own axis rather than derived from the tier — it changes what the spawned agent is asked to demonstrate, not what the tier score was. One place it does move the outcome: a quick-scored task with any bar trigger fired is spawned at the standard tier instead.
Ladder
A profile mapping each tier to a model and an effort level. Three of them; opus-centric is the default.
Concept file
One Markdown file in knowledge/ describing a single invariant, contract, or domain idea.
Drift guard
A commit-time check comparing a library bundle's pinned version against the dependency your project actually declares.
Guard
A deterministic Node script wired to a hook, blocking a tool call or a commit. Fails closed.

Getting unstuck

  • A gate fired and you don't know why → docs/GATES.md
  • Work routed to a tier you didn't expect → check which verification-bar trigger fired
  • An agent wrote a stale library API → distil a pinned bundle for it
  • Sessions keep re-deriving the same context → that is what knowledge/ is for
BigIn Skills handbook · generated with the workshop-handbook skill · 2026