Spec 001 — Superpowers-backed development workflow
Status: implemented Branch: 001-workflow-tooling
Problem
CLAUDE.md defines six workflow steps and a definition of done, but it names no technique for carrying any of them out. How a spec gets elicited, how finely tasks are sliced, whether a test is written before the code, and what "verified" means are all left to the agent to improvise, and they come out differently every session. The workflow is the point of this project, so an unreproducible workflow is the one defect that matters.
The superpowers plugin supplies exactly the missing layer — techniques for brainstorming, planning, TDD, debugging, review, and branch completion — but files its artifacts under date-named paths (docs/superpowers/specs/YYYY-MM-DD-<topic>-design.md) that ignore specs/NNN-<slug>/. Adopted naively, it would fork the project into two competing filing systems.
The workflow this spec installs
Diamonds are decisions the agent makes; hexagons are hard stops that only the human can release. The agent never sets a status field.
Step summaries
0 · Scaffold. Input: a slug. Output: specs/NNN-<slug>/ containing the three templates with placeholders substituted, and branch NNN-<slug> checked out. One command, so numbering, filing, and branch name cannot drift apart. This is the only step with no superpowers skill behind it — it is bookkeeping this project invented and superpowers has no opinion on.
1 · Spec. Technique: brainstorming. Input: a rough idea. Output: spec.md. The skill asks one question at a time, prefers multiple choice, proposes two or three alternatives before settling, and applies YAGNI ruthlessly. Its default output is a loose design document; ours must instead take the shape of specs/_template/spec.md — Given-When-Then acceptance scenarios, measurable technology-agnostic success criteria, an empty traceability table. Anything the human must decide but hasn't is written as [NEEDS CLARIFICATION: …] rather than assumed, and blocks step 2. The spec says what and why; a technology named here is a bug.
1.5 · Discovery. Technique: this project's own, fanning out with dispatching-parallel-agents when the questions are independent. Input: a spec draft with no unresolved [NEEDS CLARIFICATION]. Output: research.md. Every load-bearing claim the spec rests on is confirmed before the human is asked to approve it — against the codebase and architecture docs for internal claims, and against the outside world for external ones. The two markers divide by who can answer: [NEEDS CLARIFICATION] is answered by the human, [NEEDS VERIFICATION] is answered by reality, and neither is ever answered by assumption. Each entry in research.md records the question, the verdict, the evidence that settles it — the file and line, the endpoint response, the measured timing — and the date. A number in a success criterion is a guess until it cites a measurement here.
Answering these questions is allowed to involve throwaway spike code: written outside the repo, run once, its finding recorded, then discarded. This is the second exception to "no code without a spec and an approved plan", and it is narrow — a spike that survives into the working tree has become implementation, and implementation is gated. A refuted claim sends the spec back to step 1 rather than being patched in place, because a spec built on a false premise usually has more wrong with it than the one line that failed. Findings that outlive the feature graduate into docs/architecture/ — that is how the GW2 rate-limit budget got there, and it is what stops the next spec re-discovering it.
2 · Plan. Technique: writing-plans, header half. Input: an approved spec.md. Output:plan.md, alone. Chosen design, rejected alternatives with reasons, file structure with each file's responsibility, Global Constraints copied verbatim from the architecture docs, test strategy, risks. This is where technology enters. The plan is a proposal until the human approves it, and producing tasks.md in the same breath is the specific mistake R12 exists to prevent.
3 · Tasks. Technique: writing-plans, task half. Input: an approved plan.md. Output:tasks.md. Decomposition into tasks that each carry their own test cycle and are worth a fresh reviewer's gate — split only where a reviewer could reject one task while approving its neighbour. Within a task, each step is one action of two to five minutes: write the failing test, watch it fail, write minimal code, watch it pass, commit. Every task names the spec criterion it satisfies.
3.5 · Isolate. Technique: using-git-worktrees. An isolated worktree so a long implementation run cannot disturb the working tree, and so parallel agents do not collide.
4 · Implement. Techniques: subagent-driven-development (implementer, then a two-stage review — spec compliance, then code quality), with test-driven-development inside every task and systematic-debugging whenever something is not understood. The iron law: no production code before a failing test that demands it. Code written before its test is deleted, not retrofitted. When a diagnosis is needed, four-phase root cause analysis comes before any fix.
5 · Verify. Techniques: verification-before-completion, then requesting-code-review and receiving-code-review. Typecheck and tests clean; every acceptance scenario and success criterion traced to a named test in the spec's table; then the human reads the diff. The human's review is the gate — an agent reporting success is evidence, not proof.
6 · Capture. Technique: finishing-a-development-branch. Merge or PR, worktree cleanup, spec status set to implemented, and any decision that outlived the feature moved into a spec or docs/. A decision that exists only in chat history is lost.
Bug fixes skip steps 0 to 3 entirely: systematic-debugging to find the root cause, then a failing regression test, then the fix, then step 5. No spec directory. The exception exists so small fixes stay cheap — it is not a route around TDD or review.
User stories
Ordered by priority. Each story must be independently testable and shippable — if only P1 ships, there is still something usable.
P1 — Each workflow step has one named technique
As the human running this project, I want every step of the workflow to be carried out by a specific, named superpowers skill, so that two runs of the same step produce comparably structured work instead of whatever the agent felt like doing that day.
Independent test: with the plugin installed and CLAUDE.md updated, start a fresh session and ask for a new feature; the agent announces brainstorming before asking anything, and does not begin planning until the spec is written and approved.
Acceptance scenarios
- Given a fresh session on this repo, when the human asks to build a new feature, then the agent invokes
superpowers:brainstormingbefore any other action, including clarifying questions. - Given an approved
spec.md, when the human asks for a plan, then the agent invokessuperpowers:writing-plansand writes tospecs/NNN-<slug>/plan.mdandtasks.md, not todocs/superpowers/. - Given an approved
tasks.md, when implementation starts, then the agent usessuperpowers:subagent-driven-developmentand each task is implemented test-first. - Given a superpowers skill instruction that contradicts
CLAUDE.md, when the agent must choose, thenCLAUDE.mdwins and the agent says which instruction it is overriding. - Given a
spec.mdcontaining an unresolved[NEEDS CLARIFICATION: …]marker, when the human asks for a plan, then the agent refuses and names the unresolved markers. - Given a
spec.mdwhose status is notapproved, when the human asks for a plan, then the agent refuses, and asks for approval of the spec instead. - Given a
plan.mdwhose status is notapproved, when the human asks for tasks or for implementation, then the agent refuses, and asks for approval of the plan instead. - Given an approved spec, when the agent produces the plan, then it produces
plan.mdonly —tasks.mdstays untouched until the plan is approved in turn.
P2 — Plan artifacts are directly executable by superpowers
As the agent, I want plan.md and tasks.md to match the structure superpowers' executors expect, so that executing-plans and subagent-driven-development run against our files unmodified.
Independent test: point superpowers:executing-plans at a filled-in tasks.md from the template; it loads the file, finds checkbox steps, and needs no reformatting to proceed.
Acceptance scenarios
- Given the plan template, when
writing-plansfills it in, thenplan.mdcarries a Global Constraints section holding the project-wide rules that every task inherits. - Given the tasks template, when
writing-plansfills it in, thentasks.mdcarries the agentic-worker header naming the required execution sub-skill, and every step is a- [ ]checkbox. - Given a filled-in
tasks.md, when a task is read in isolation, then it names the spec criterion it satisfies and the test that verifies it.
P3 — Starting a feature is one command
As the human, I want a single command to open a new feature, so that numbering, templating, and branching cannot drift apart.
Independent test: run the command on a repo whose highest spec is 001; it creates specs/002-<slug>/ from the template with all three files and switches to branch 002-<slug>.
Acceptance scenarios
- Given existing specs
001-…, when the human runs the scaffold command with a slug, thenspecs/002-<slug>/exists withspec.md,plan.md, andtasks.mdcopied from the template and theirNNN/slug placeholders substituted. - Given the same command, when it completes, then the working branch is
002-<slug>.
Superseded in part by spec 009 (2026-08-03). Step 0 no longer runs as one shared-tree command: it allocates the number, creates an isolated worktree via
EnterWorktree, and scaffolds inside it. Scenario 2 is intentionally reversed — the shared tree never changes branch (009 SC2); the worktree carriesNNN-<slug>instead. The numbered-directory guarantee (scenario 1) still holds. The traceability rows below point to the tests as they stand after 009.
P4 — Bug fixes bypass the spec, not the discipline
As the human, I want a bug fix to skip the spec directory but still be forced through diagnosis and a failing regression test, so that the constitution's exception does not become a hole.
Acceptance scenarios
- Given a reported bug, when the human asks for a fix, then the agent invokes
superpowers:systematic-debuggingbefore proposing a change, and nospecs/NNN-…/directory is created. - Given an identified root cause, when the fix is written, then a failing regression test is committed before the code that makes it pass.
P5 — A spec is confirmed before it is approved
As the human, I want every claim a spec depends on to have been checked against the codebase and the outside world before I am asked to approve it, so that the plan step discovers design trade-offs rather than discovering the feature was never possible.
Independent test: draft a spec asserting something false about this repo — a module that does not exist, or a GW2 API endpoint that returns a field it does not return — and ask for approval; the agent runs discovery, records the refutation in research.md, and sends the spec back to step 1 instead of carrying the claim forward.
Acceptance scenarios
- Given a spec draft with no unresolved
[NEEDS CLARIFICATION], when the human asks to approve it, then the agent first producesresearch.mdwith a verdict and cited evidence for every[NEEDS VERIFICATION]marker in the spec. - Given a
spec.mdcontaining an unresolved[NEEDS VERIFICATION: …]marker, when the human asks for a plan, then the agent refuses and names the unverified markers, exactly as it does for[NEEDS CLARIFICATION]. - Given a discovery question about this repo's own code, when it is answered in
research.md, then the entry cites a file path and line rather than a recollection. - Given a claim that discovery refutes, when the finding is recorded, then the spec returns to step 1 for rewrite and its status stays at
draft. - Given a success criterion stating a number, when the spec reaches the approval gate, then that number cites a measurement recorded in
research.md. - Given a discovery finding that outlives the feature, when step 6 runs, then it has been moved into
docs/architecture/rather than left only in the spec directory.
Requirements
- R1 — The superpowers plugin is declared in project-scoped, version-controlled settings so that a fresh clone of this repo gets the same workflow without manual setup.
- R2 —
CLAUDE.mdmaps every workflow step to the superpowers skill that carries it out. - R3 —
CLAUDE.mddeclares the spec and plan path overrides that redirectbrainstormingandwriting-plansintospecs/NNN-<slug>/. - R4 —
CLAUDE.mdstates the precedence rule: superpowers skills are mandatory within a step; where a skill and this constitution conflict, the constitution wins. - R5 —
spec.md's existing shape (Given-When-Then scenarios, measurable technology-agnostic success criteria,[NEEDS CLARIFICATION]markers, traceability table) is the required output shape forbrainstorming, overriding its looser default. - R6 —
plan.mdcarries superpowers' plan header fields — Goal, Architecture, Tech Stack, Global Constraints, File Structure — alongside the existing Alternatives, Test strategy, and Risks sections. - R7 —
tasks.mdcarries superpowers' agentic-worker header and task structure, with steps as- [ ]checkboxes sized at one action each. - R8 — Global Constraints are populated from
docs/architecture/typescript.mdandstack.md, with values copied verbatim rather than summarised. - R9 — A scaffold command allocates the next spec number, copies the template, substitutes placeholders, and creates the matching branch.
- R10 —
CLAUDE.mdroutes bug fixes tosystematic-debuggingthentest-driven-development, with no spec directory. - R11 — Each artifact template carries a
Status:field whose default value is the pre-approval one, so an unapproved artifact is unapproved by construction rather than by the agent remembering. - R12 —
CLAUDE.mdstates the artifact gates as hard stops, in the human's words: the only way to produce a plan is for the spec's status to beapprovedby the human; the only way to produce tasks is for the plan's status to beapprovedby the human. Each artifact is produced alone, then handed over for approval. Status is set by the human, never by the agent. - R13 —
CLAUDE.mdcarries the workflow diagram and the one-line-per-step bindings, and links to this spec for the long-form step summaries rather than duplicating them. The constitution stays short enough to be read every session. - R14 — The conformance of the artifacts in R1–R13 is checked by an automated test suite, so the definition of done applies to this spec as it will to every later one.
- R15 —
specs/_template/carries aresearch.mdwhose shape is one entry per question, each with the question, a verdict of confirmed or refuted, the evidence that settles it, and the date. The scaffold command of R9 copies it alongside the other three artifacts. - R16 —
CLAUDE.mdstates the second gate condition in the same hard-stop terms as R12: a spec whose[NEEDS VERIFICATION]markers are not all resolved inresearch.mdcannot be approved, and an unapproved spec cannot become a plan. It also names the division between the two markers — the human answers clarification, reality answers verification. - R17 —
CLAUDE.mdnames throwaway discovery spikes as the second exception to "no code without a spec and an approved plan", bounded to code that is run outside the repo and discarded, with the finding recorded inresearch.md.
Success criteria
Measurable and technology-agnostic — outcomes, not implementation.
- SC1 — A fresh clone plus one dependency install yields a session in which all fourteen superpowers skills listed in the appendix are available, with no per-machine configuration step.
- SC2 — Every step in the workflow diagram names at least one technique; zero steps are left to agent discretion. The sole exception is step 0, which is this project's own bookkeeping.
- SC3 — No workflow artifact is written outside
specs/NNN-<slug>/— the count of files under any date-named superpowers default path stays at zero. - SC4 — Opening a new feature takes one command and produces a correctly numbered directory, substituted placeholders, and a matching branch.
- SC5 — Every acceptance scenario and success criterion in this spec maps to a named automated test, and the suite passes.
- SC6 — Typecheck passes with no
anyand no unexplained escape hatches in anything added here. - SC7 — Across every spec directory in the repo, the count of filled-in plans whose spec is not
approvedis zero, and the count of filled-in task lists whose plan is notapprovedis zero. - SC8 — Across every spec directory in the repo, the count of
approvedspecs carrying an unresolved[NEEDS VERIFICATION]marker is zero, and the count ofapprovedspecs with noresearch.mdentries is zero.
Out of scope
- CI wiring (GitHub Actions running lint/typecheck/test). Deferred to its own spec; this one only makes the suite exist and pass locally.
- Authoring skills of our own.
writing-skillsis available if a gap shows up later; we do not pre-emptively invent skills superpowers already covers. - Rewriting
docs/project-brief.mdor the architecture docs. They become the source for Global Constraints; their content is untouched. - Any GW2 domain code. This spec builds the workbench, not the product.
- Pinning a superpowers version. We track the marketplace's current release and revisit if it breaks.
Assumptions
This spec was the first subject of its own step 1.5. Every claim below that describes the outside world was confirmed on 2026-07-23 against superpowers v6.1.1 — see research.md for the evidence.
- Superpowers' documented extension points hold: both
brainstormingandwriting-plansstate that user preferences override their default artifact locations, andusing-superpowersstates thatCLAUDE.mdtakes precedence over skills. The integration rests on these two behaviours. Confirmed, V1 — all three statements verbatim in the shipped skill files. The behavioural half, that an agent honours them in a live session, is not statically verifiable and stays a review-time check. - The plugin ships all fourteen skills named in the appendix, under those names. Confirmed, V2 — exact one-to-one match, no renames.
- Vitest is the project's test runner, already fixed by
docs/architecture/stack.md; this spec is its first consumer and may add it as a dev dependency. - The plugin ships a
SessionStarthook that runs on startup, clear, and compact. Confirmed, V3 — matcher is exactlystartup|clear|compact,async: false. Cost measured at ~800 tokens per fire, which includes every compaction, not just session start. Accepted with the number attached. - The plugin is not yet installed; everything above was verified against the upstream repository. Installing it is implementation work under R1. F5.
- Skill invocation is a behavioural expectation of the agent, not a mechanically enforceable one. Tests here verify that the instructions and artifacts are in place, which is the furthest automation can reach; the behaviour itself is verified by the human in review.
Appendix — skill inventory
All fourteen skills the plugin ships, and where each lands. Nothing is disabled; the "on demand" skills are available without being bound to a step.
| Skill | Where it fires |
|---|---|
using-superpowers | Always. Establishes that skills are checked before any response, and defers to CLAUDE.md. |
brainstorming | Step 1. Elicits the spec. |
writing-plans | Steps 2 and 3, split across the approval gate. |
executing-plans | Step 4, alternative to subagent-driven-development for small features. |
subagent-driven-development | Step 4, default. Implementer plus two-stage review per task. |
test-driven-development | Step 4, inside every task, and in the bug-fix route. |
systematic-debugging | Step 4 on demand; first step of the bug-fix route. |
using-git-worktrees | Step 3.5. |
verification-before-completion | Step 5, before anything is called done. |
requesting-code-review | Step 5. |
receiving-code-review | Step 5, on what the review returns. |
finishing-a-development-branch | Step 6. |
dispatching-parallel-agents | Step 1.5, when discovery questions are independent; on demand elsewhere. |
writing-skills | On demand, if a gap appears that superpowers does not cover. |
Traceability
Each acceptance scenario and success criterion maps to a named test. SC5's own tests assert this table has no empty cells and that every test it names exists in the suite, so it cannot rot into decoration.
Read the middle column before trusting the right one. A proxy test asserts that the instruction is present and unambiguous in CLAUDE.md — it does not observe an agent obeying it, which no test file can. Those criteria are verified by the human at review time; the test is a floor, not a proof.
| Criterion | Kind | Test |
|---|---|---|
| P1 #1 | proxy + human | P1 #1 (proxy): a new feature starts with brainstorming |
| P1 #2 | proxy + human | P1 #2 (proxy): plan artifacts are directed into the spec directory |
| P1 #3 | proxy + human | P1 #3 (proxy): implementation is subagent-driven and test-first |
| P1 #4 | proxy + human | P1 #4 (proxy): a conflict must be surfaced, not silently resolved |
| P1 #5 | proxy + human | P1 #5–#7 (proxy): an unmet gate is refused, and the reason is named |
| P1 #6 | proxy + human | P1 #5–#7 (proxy): an unmet gate is refused, and the reason is named |
| P1 #7 | proxy + human | P1 #5–#7 (proxy): an unmet gate is refused, and the reason is named |
| P1 #8 | artifact | R12: each artifact is produced alone |
| P2 #1 | artifact | P2 #1: the plan template carries a Global Constraints section |
| P2 #2 | artifact | P2 #2: the tasks template names the execution sub-skill and uses checkbox steps |
| P2 #3 | artifact | P2 #3: a task names the spec criterion it satisfies and the test that verifies it |
| P3 #1 | integration | SC4 / R4: writeScaffold writes the four artifacts with placeholders substituted, R3 / SC7: allocate numbers from origin/main and mutates nothing |
| P3 #2 | superseded by 009 | branch now owned by EnterWorktree, not the scaffold (human-verified); the shared tree never switches — see spec 009 SC2/SC3 |
| P4 #1 | proxy + human | P4 #1–#2 (proxy): a bug fix creates no spec directory but keeps the discipline |
| P4 #2 | proxy + human | R10: CLAUDE.md routes bug fixes to debugging then a failing regression test |
| P5 #1 | proxy + human | P5 #1 (proxy): discovery produces research.md before the approval gate |
| P5 #2 | artifact | SC8: no approved spec carries an unresolved [NEEDS VERIFICATION], and no approved spec lacks research entries |
| P5 #3 | proxy | P5 #3 (proxy): the research template demands cited evidence, not recollection |
| P5 #4 | proxy + human | P5 #4 (proxy): a refuted claim returns the spec to step 1 |
| P5 #5 | proxy | P5 #5 (proxy): a success criterion stating a number must cite a measurement |
| P5 #6 | proxy | P5 #6 (proxy): the research template names what should graduate to docs/ |
| SC1 | known-unsatisfiable | SC1: a fresh clone plus one dependency install yields all fourteen skills — see research.md V6. Ships as an it.fails() case: it passes by failing, and turns red if the platform ever makes SC1 true. |
| SC2 | artifact | SC2: every workflow step names at least one technique |
| SC3 | repo scan | SC3: no workflow artifact is written outside specs/NNN-<slug>/ |
| SC4 | integration | SC4 / R4: writeScaffold writes the four artifacts with placeholders substituted — the correctly-numbered directory; "one command" superseded by 009 (now allocate → EnterWorktree → scaffold) |
| SC5 | repo scan | SC5: the traceability table has no empty cells, SC5: every test the table names exists in the suite |
| SC6 | repo scan | SC6: no any and no unexplained escape hatches |
| SC7 | repo scan | SC7: no filled-in plan whose spec is unapproved, and no filled-in task list whose plan is unapproved |
| SC8 | repo scan | SC8: no approved spec carries an unresolved [NEEDS VERIFICATION], and no approved spec lacks research entries |