Plan 001 — Superpowers-backed development workflow
Status: approved Written in plan mode from spec.md (approved) and research.md. Approved by the human before any code is written.
Goal
Make the six-step workflow in CLAUDE.md reproducible by binding every step to a named superpowers technique, filing every artifact under specs/NNN-<slug>/, and proving the whole arrangement with a test suite that fails when the arrangement drifts.
Approach
The feature is four thin layers plus a suite that holds them to account. None of it is product code, and nothing here touches apps/ or packages/ — those directories still do not exist.
1 · Configuration. A new version-controlled .claude/settings.json enabling the plugin from Anthropic's official marketplace, which is already registered on every machine:
{ "enabledPlugins": { "superpowers@claude-plugins-official": true } }Amended during T2 (human, 2026-07-23). This section previously specified a pinned custom marketplace. That was attempted, failed to register from project settings, and pinning was then dropped by decision — see research.md V7. The version now floats and the fourteen-name inventory test is the tripwire for drift. settings.local.json stays as it is: machine-local permissions, not workflow configuration.
2 · Instruction. CLAUDE.md gains a Techniques section: a bindings table of one line per step, the two path overrides that redirect brainstorming and writing-plans into specs/NNN-<slug>/, the precedence rule, the two gates stated as hard stops, and the two exceptions (bug fixes, discovery spikes). Long-form step summaries stay in this spec and are linked, not copied — R13 exists because the constitution is read every session and a constitution nobody finishes reading is not one.
3 · Artifacts. specs/_template/ grows a fourth file, research.md, and the existing three are reshaped: plan.md takes superpowers' header fields (Goal, Architecture, Tech Stack, Global Constraints, File Structure) alongside the sections it already has, tasks.md takes the agentic-worker header and - [ ] steps, and every Status: line defaults to its pre-approval value rather than offering a menu. That last point is R11 and it is load-bearing: a template reading draft | approved | implemented is a prompt to pick one, and the safe value should be the one you get by doing nothing.
4 · Automation. scripts/new-spec.ts allocates the next number, copies all four templates, substitutes placeholders, and creates the branch. Split into a pure core and a thin shell: number allocation and placeholder substitution are pure functions over strings and directory listings, so they unit-test without touching git; the filesystem and git checkout -b calls sit in one entry point tested once against a temporary repository. A /new-spec slash command wraps it for ergonomics but holds no logic, because logic in a markdown command cannot be tested.
5 · Conformance. tests/workflow/ under Vitest, asserting the four layers above are in place. This suite is the first consumer of Vitest in the repo.
The honest boundary, which the spec already concedes and this plan does not try to talk its way around: these tests verify that instructions and artifacts exist, never that an agent obeyed them. "The agent invokes brainstorming first" is not assertable from a test file. Every such scenario gets an artifact-level proxy test — the instruction is present, unambiguous, and in the right file — and is marked in the traceability table as human-verified. Proxy tests are labelled as proxies so nobody later mistakes a green suite for proof of behaviour.
Architecture
.claude/settings.json → enables the plugin (layer 1)
CLAUDE.md → binds steps to techniques (layer 2)
specs/_template/*.md → artifact shapes (layer 3)
scripts/new-spec.ts → scaffold: pure core + shell (layer 4)
tests/workflow/*.test.ts → holds 1-4 to account (layer 5)Dependencies run one way: tests read the other four, the other four know nothing about the tests.
Tech stack
- Node ≥ 22.18, pnpm 11.15.1 — as
package.jsonalready declares, with the floor raised (see below). - Vitest — mandated by
docs/architecture/stack.md; this spec is its first consumer. - TypeScript for the scaffold and the tests. Verified during planning: Node 26.5.0 executes
.tsfiles directly with unflagged type stripping, sonode scripts/new-spec.tsruns with no new runtime dependency — notsx, no build step. This is whyengines.nodemoves from>=22to>=22.18, the version where stripping became unflagged. TypeScript itself is added as a dev dependency for typechecking only.
New dev dependencies, justified here per docs/architecture/typescript.md: vitest (mandated by the stack doc), typescript (needed for pnpm typecheck, which the definition of done requires), and @types/node (the scaffold touches node:fs and node:child_process).
Global Constraints
Copied verbatim from the architecture docs per R8. Every task inherits these.
From docs/architecture/typescript.md:
- No
any. Not in app code, not in tests. Useunknownplus narrowing, or model the type properly. If a third-party type forces it, isolate it behind one typed adapter and comment why. - No non-null assertions (
!) to silence the compiler. - No
@ts-expect-errorwithout a comment explaining what is expected and when it can be removed. - Validate everything crossing a boundary (GW2 API responses, HTTP input) at runtime, not just at the type level.
- Prefer pure functions for domain logic. The optimizer must be testable without a network or a database.
- Match the style of surrounding code. No new dependency without justification in the spec or plan.
From docs/architecture/stack.md:
- Monorepo, pnpm workspaces.
apps/api— NestJS (TypeScript).apps/web— React (TypeScript).packages/*— shared code when sharing is real, not speculative. - Vitest everywhere, both apps.
From CLAUDE.md: typecheck clean, tests pass, every acceptance scenario and success criterion covered by a test whose name traces to it, no unexplained escape hatches, human reviews the diff.
File Structure
| Path | Change | Responsibility |
|---|---|---|
.claude/settings.json | new | Enables superpowers@claude-plugins-official. Version-controlled, team-shared. |
CLAUDE.md | modified | Techniques table, path overrides, precedence rule, both gates, both exceptions. Links to spec 001 for long form. |
specs/_template/spec.md | modified | Status: draft by default; [NEEDS VERIFICATION] documented alongside [NEEDS CLARIFICATION]. |
specs/_template/plan.md | modified | Status: proposed by default; gains Goal, Architecture, Tech Stack, Global Constraints, File Structure. |
specs/_template/tasks.md | modified | Agentic-worker header naming the execution sub-skill; every step a - [ ] checkbox; each task names its spec criterion. |
specs/_template/research.md | new | Step 1.5 artifact: question, verdict, evidence, date. |
scripts/new-spec.ts | new | Scaffold. Pure core (nextSpecNumber, substitutePlaceholders) + main() doing fs and git. |
.claude/commands/new-spec.md | new | Thin /new-spec wrapper. No logic. |
tests/workflow/settings.test.ts | new | Layer 1: R1, and SC1 as an it.fails() known-unsatisfiable case. |
README.md | modified | Setup: the two commands a fresh clone needs. |
tests/workflow/constitution.test.ts | new | Layer 2: R2, R3, R4, R10, R13, R16, R17, SC2. |
tests/workflow/templates.test.ts | new | Layer 3: R5, R6, R7, R11, R15, and R8's verbatim check. |
tests/workflow/scaffold.test.ts | new | Layer 4: R9, P3, SC4. |
tests/workflow/repo-invariants.test.ts | new | Cross-cutting: SC3, SC5, SC7, SC8. |
package.json | modified | Dev deps; test and typecheck scripts; engines.node → >=22.18. |
tsconfig.json | new | Strict. noEmit — typechecking only, Node does the stripping. |
Data & contracts
Only the scaffold has a contract worth naming:
type ScaffoldRequest = { slug: string; specsDir: string };
type ScaffoldResult = { number: string; dir: string; branch: string; files: string[] };
function nextSpecNumber(existing: readonly string[]): string; // ["001-a"] -> "002"
function substitutePlaceholders(body: string, n: string, slug: string): string;nextSpecNumber takes directory names, not a path, which is what makes it pure and what lets the allocation logic be tested without a fixture tree.
Everything else in this feature is markdown and JSON with no runtime contract.
Test strategy
Vitest at the repo root, tests/workflow/. Fixtures are temporary directories under os.tmpdir(); the scaffold's git test initialises a throwaway repo and asserts the branch it lands on. No network, no mocking of the filesystem — the suite reads the real committed files, because a test that reads a mock of CLAUDE.md proves nothing about CLAUDE.md.
Test names carry their criterion so the traceability table is mechanical: R2: CLAUDE.md binds every workflow step to a named technique, P3 #1: scaffold creates the next numbered directory.
Three kinds of test, and it matters which is which:
- Artifact tests — read a committed file, assert its structure. Most of the suite.
- Proxy tests — for behavioural scenarios (P1 #1, #3, parts of #2 and P4 #1). Assert the instruction is present and unambiguous. Named with a
(proxy)suffix and marked human-verified in the table. - Integration tests — the scaffold against a temp repo.
Notable cases:
- SC5 is self-referential and worth building carefully. A test parses the traceability table, asserts no cell is empty, and asserts every test name it references actually exists in the suite. This is what stops the table rotting into decoration.
- SC2 parses the bindings table in
CLAUDE.mdand asserts every step names at least one technique, with step 0 the sole allowed exception — encoded as an explicit allowlist of one, so a second unbound step fails. - SC3, SC7, SC8 scan
specs/*/anddocs/superpowers/. They are repo-wide invariants, not spec-001 assertions: they must keep passing for every future spec, which is the point. - R8 asserts the Global Constraints block matches the architecture docs verbatim, by string containment. Summarising them is the failure mode, so the test forbids it.
- SC6 is
pnpm typecheckplus a grep foranyin added files, run in CI later — out of scope here per the spec, so locally for now.
Deliberately not tested: that an agent actually invokes a skill; that superpowers' internals behave as documented (discovery covered that, and it is upstream's suite to own).
Alternatives considered
- Scaffold as a pure markdown slash command — rejected. R14 requires the scaffold be covered by automated tests, and a markdown command's numbering logic cannot be asserted. The command remains as a wrapper.
tsxor a build step for the scaffold — rejected. Node runs.tsdirectly (verified, Node 26.5.0), so a dependency would buy nothing but a floor bump we need anyway.- Scaffold in plain
.mjsto dodge the toolchain — rejected. It would be the only untypechecked file in the repo, against the constitution's own rules, to save one dev dependency we need fortypecheck. - Copying superpowers' skills into the repo and editing them — rejected. Forks upstream, and discovery confirmed the documented override points make it unnecessary.
- Putting the long-form step summaries in
CLAUDE.md— rejected by R13. The constitution is read every session; spec 001 is read when someone wants the detail. - Vitest workspace config now — rejected as premature.
apps/andpackages/do not exist. One root config, split when there is something to split.
Risks
- SC1 ships known-unsatisfiable — decided, not open. A fresh clone needs two commands, not one (
research.mdV6, confirmed by spike). The human chose to leave SC1 as written rather than amend it, so the spec carries a criterion that cannot pass. Mitigation is disclosure, not repair: anit.fails()test keeps the suite green while recording the gap, the traceability table names it, and the test flips red if the platform ever makes SC1 true. The residual risk is that "all criteria have passing tests" becomes a slightly weaker claim for this spec than for later ones — anyone reading the suite should know SC1 is the exception, which is why it is stated here and in the table rather than only in a commit message. - Behavioural rules are unenforceable. A green suite means the instructions are in place, not that they were followed. Mitigation: proxy tests are labelled, the table marks them human-verified, and the first real feature through the workflow is the actual test.
- Upstream drift. A superpowers release could rename a skill or move a default path, silently breaking a binding. This is sharper than it first looked: the official marketplace auto-updates by default, so drift arrives on its own timetable rather than when we ask for it. Pinning was tried and dropped (
research.mdV7), so the only mitigation is the fourteen-name inventory test: it turns a silent rename into a failing suite. That is a detection mechanism, not a prevention one, and it is the accepted position. The upside is that this matches the spec's Out of scope line exactly — no contradiction to absorb. - Node floor. Raising
engines.nodeto>=22.18will break anyone on 22.0–22.17. Acceptable: single developer, running 26.5.0. - Template/test drift. Tests asserting template structure must be updated with the templates. Mitigation: both live in this spec's diff, and the SC5 test catches an untraced criterion.
Open questions
Does the plugin trust prompt violate SC1?Resolved — SC1 is unachievable as worded. Seeresearch.mdV6. A spike outside the repo confirmed that a project.claude/settings.jsonenabling the plugin installs nothing:claude plugin liststill reports none installed. Superpowers sources from an external URL even inside the official marketplace, so the v2.1.195 install step applies, and the installed CLI (2.1.218) is past that threshold. A fresh clone needs two commands, not one.Decision (human, 2026-07-23): leave SC1 as written and accept that it fails. Amending it was offered and declined. The spec text stands unchanged and its status stays
approved; no gate reopens.This plan therefore carries a known-unsatisfiable success criterion, and the job is to make that visible rather than quiet. SC1 gets a real test, written with Vitest's
it.fails():it.fails('SC1: a fresh clone plus one dependency install yields all fourteen skills (known-unsatisfiable — see research.md V6)', …)it.fails()passes when its body fails, so the suite stays green while encoding the gap honestly, and SC5's no-empty-cells guarantee still holds because SC1 has a named test like every other row. The useful side effect: if the platform ever stops requiring the second command, the test stops failing,it.fails()turns red, and you are told to revisit SC1 — the gap becomes a tripwire instead of a comment nobody rereads.R1's implementation is unaffected either way, and improves:
--scope projectwrites theenabledPluginsentry itself, so.claude/settings.jsoncan be generated by the install command and committed rather than hand-authored. Setup becomes two commands, documented inREADME.md.Should
/new-specalso open the branch? — decided (human, 2026-07-23): yes. Starting a new spec creates its branch. R9 stands as written, the scaffold mutates git state, andscaffold.test.tsasserts the resulting branch in a temporary repository. No longer open.
Nothing else is open. The pinning mechanism is settled as a design choice below and its one unverified detail is carried as research.md V7, to be confirmed by the task that writes the settings file rather than by a decision here.
Traceability
Populated during implementation. The SC5 test asserts this table has no empty cells and that every test it names exists, so it cannot be left half-filled.