Skip to content

Tasks 001 — Superpowers-backed development workflow ​

Execution skill: superpowers:subagent-driven-development — one implementer per task, then a two-stage review (spec compliance, then code quality). superpowers:test-driven-development applies inside every task: no production code before a failing test that demands it. Reach for superpowers:systematic-debugging on any surprise rather than guessing.

Derived from plan.md (approved). Each task is small, independently verifiable, and reviewed as its own diff — the split points are where a reviewer could reject one task while approving its neighbour. A task is done only when it satisfies the definition of done in CLAUDE.md.

Global Constraints in plan.md apply to every task and are not repeated per task.


T1 — Test harness, and the first invariant it proves ​

Satisfies: SC3, and the tooling half of SC5/SC6. Why first: every later task is TDD, and TDD needs somewhere to put a failing test.

  • [x] Add vitest, typescript, @types/node as dev dependencies; justify each in the commit body (Global Constraints require new dependencies be justified — plan.md Tech stack has the wording).
  • [x] Add tsconfig.json: strict, noEmit, module resolution for Node.
  • [x] Add test and typecheck scripts to package.json.
  • [x] Raise engines.node from >=22 to >=22.18 — the floor where unflagged type stripping landed.
  • [x] RED: write tests/workflow/repo-invariants.test.ts with SC3: no workflow artifact is written outside specs/NNN-<slug>/, asserting zero files under docs/superpowers/. Watch it fail for the right reason — harness absent, not assertion wrong.
  • [x] GREEN: make pnpm test run. The assertion should already hold; if it does not, that is a finding.
  • [x] Confirm the assertion is not vacuous: plant a file under docs/superpowers/, watch SC3 fail, remove it.
  • [x] Commit.

Verified by: pnpm test green, pnpm typecheck clean, SC3: no workflow artifact is written outside specs/NNN-<slug>/.


T2 — Plugin configuration, pinned ​

Satisfies: R1, SC1. Note: deliberately early so the remaining tasks can dogfood the skills they are installing.

  • [x] Resolve research.md V7 first — it is open, and this task is the one that closes it. Register obra/superpowers as a marketplace, install, and record in research.md: whether it registers cleanly, the plugin id it yields, and whether autoUpdate really defaults to false.
  • [x] If V7 is refuted, stop and report rather than improvising. Refuted: the inline marketplace never registered from project settings. Stopped; the human then dropped pinning entirely.
  • [x] RED: tests/workflow/settings.test.ts — R1: the plugin is enabled in version-controlled project settings.
  • [x] GREEN: write .claude/settings.json enabling superpowers@claude-plugins-official.
  • [x] Add the it.fails() case for SC1, named SC1: a fresh clone plus one dependency install yields all fourteen skills (known-unsatisfiable — see research.md V6). Assert the real thing — that one install command suffices — so the test flips red if the platform ever makes SC1 true.
  • [x] Install the plugin and confirm: v6.1.1, enabled, fourteen SKILL.md files on disk.
  • [x] Document the two setup commands in README.md: pnpm install, then the plugin install.
  • [x] Commit.

Verified by: R1: the plugin is enabled in version-controlled project settings; SC1: … (known-unsatisfiable) passing as an it.fails(); research.md V7 closed with evidence.


T3 — Template: research.md, and pre-approval status defaults ​

Satisfies: R11, R15, and R5's marker half.

  • [x] RED: tests/workflow/templates.test.ts — R15: the template carries research.md with the step 1.5 shape and R11: every template's Status defaults to its pre-approval value.
  • [x] GREEN: add specs/_template/research.md — one entry per question with question, verdict, evidence, date. Model it on this spec's own research.md, which is the worked example.
  • [x] GREEN: change each template's Status: line from a menu to the single safe default — draft for spec, open for research, proposed for plan. A menu is an invitation to pick; the default should be the safe one.
  • [x] GREEN: document [NEEDS VERIFICATION] alongside [NEEDS CLARIFICATION] in specs/_template/spec.md, naming the split — the human answers one, reality answers the other.
  • [x] Commit.

Verified by: R15: the template carries research.md with the step 1.5 shape, R11: every template's Status defaults to its pre-approval value, R5: the spec template documents both marker kinds.


T4 — Template: plan.md header fields and Global Constraints ​

Satisfies: R6, R8, P2 #1.

  • [x] RED: R6: the plan template carries the superpowers header fields (Goal, Architecture, Tech Stack, Global Constraints, File Structure) and P2 #1: the plan template carries a Global Constraints section.
  • [x] RED: R8: Global Constraints are copied verbatim from the architecture docs — assert by string containment against docs/architecture/typescript.md and stack.md. Summarising is the failure mode; the test must forbid it.
  • [x] GREEN: add the five header sections to specs/_template/plan.md, keeping the existing Alternatives, Test strategy, Risks, and Open questions sections.
  • [x] GREEN: populate Global Constraints verbatim — all 6 rules from typescript.md, all 12 from stack.md.
  • [x] Confirm R8 catches rewording, not just absence: soften one rule, watch it fail, restore.
  • [x] Commit.

Verified by: R6: the plan template carries the superpowers header fields, R8: Global Constraints are copied verbatim from the architecture docs, P2 #1: the plan template carries a Global Constraints section.


T5 — Template: tasks.md agentic-worker shape ​

Satisfies: R7, P2 #2, P2 #3.

  • [x] RED: P2 #2: the tasks template names the execution sub-skill and uses checkbox steps and P2 #3: a task names the spec criterion it satisfies and the test that verifies it.
  • [x] GREEN: reshape specs/_template/tasks.md — execution-skill header, - [ ] steps sized at one action, a Satisfies: line and a Verified by: line per task.
  • [x] Cross-check against this file, which already follows the shape. If the template and this file disagree, one of them is wrong — reconcile rather than letting them drift. Made mechanical: one shape assertion runs over both, so they cannot drift silently.
  • [x] Confirm the shape check has teeth — strip a Satisfies: line, watch it name the offending task, restore.
  • [x] Commit.

Verified by: P2 #2: the tasks template names the execution sub-skill and uses checkbox steps, P2 #3: a task names the spec criterion it satisfies and the test that verifies it, R7: the tasks template carries the agentic-worker structure.


T6 — Constitution: technique bindings and path overrides ​

Satisfies: R2, R3, R4, R13, SC2, and proxies for P1 #1–#4.

  • [x] RED: tests/workflow/constitution.test.ts — SC2: every workflow step names at least one technique, parsing the bindings table with step 0 as an explicit allowlist of one, so a second unbound step fails.
  • [x] RED: R3: CLAUDE.md declares the spec and plan path overrides, R4: CLAUDE.md states that the constitution outranks skills.
  • [x] RED: proxy tests for P1 #1–#4, each suffixed (proxy) — they assert the instruction is present and unambiguous, never that an agent obeyed it.
  • [x] GREEN: add the Techniques section to CLAUDE.md — bindings table, the two path overrides, the precedence rule. Link to specs/001-workflow-tooling/spec.md for the long-form step summaries; do not copy them (R13).
  • [x] Add steps 0, 1.5 and 3.5 to the workflow list, which previously stopped at the original six.
  • [x] Confirm SC2 has teeth — add an unbound step to the table, watch it named in the failure, restore.
  • [x] Re-read CLAUDE.md end to end afterwards. If it no longer reads in one sitting, R13 has been violated regardless of what the tests say. 654 words, cap is 1200.
  • [x] Commit.

Verified by: SC2: every workflow step names at least one technique, R2: CLAUDE.md maps every workflow step to a technique, R3: CLAUDE.md declares the spec and plan path overrides, R4: CLAUDE.md states that the constitution outranks skills, R13: CLAUDE.md links rather than duplicates the step summaries, P1 #1–#4 (proxy).


T7 — Constitution: gates and the two exceptions ​

Satisfies: R10, R12, R16, R17, and proxies for P1 #5–#8, P4 #1–#2.

  • [x] RED: R12: CLAUDE.md states both approval gates as hard stops, R16: CLAUDE.md states the discovery gate and the marker split, R17: CLAUDE.md bounds the discovery-spike exception, R10: CLAUDE.md routes bug fixes to debugging then a failing regression test.
  • [x] RED: proxies for P1 #5–#8 and P4 #1–#2.
  • [x] Scope the gate assertion to the ## Gates section — matched against the whole file it passed before the section existed, which makes it decoration rather than a test.
  • [x] GREEN: write the gates in the human's own terms — spec approved before a plan exists, plan approved before tasks exist, each artifact produced alone and handed over, status set by the human and never by the agent.
  • [x] GREEN: write both exceptions — bug fixes (no spec directory; debugging then a failing regression test first) and discovery spikes (run outside the repo, discarded, finding recorded in research.md).
  • [x] Confirm both new rules have teeth — soften the status rule and the "discarded" clause, watch each fail, restore.
  • [x] Commit.

Verified by: the four R… tests above plus P1 #5–#8 (proxy) and P4 #1–#2 (proxy).


T8 — Repo-wide gate invariants ​

Satisfies: SC7, SC8. Note: these must hold for every future spec, not just this one. Write them to scan specs/*/.

  • [x] RED: SC7: no filled-in plan whose spec is unapproved, and no filled-in task list whose plan is unapproved.
  • [x] RED: SC8: no approved spec carries an unresolved [NEEDS VERIFICATION], and no approved spec lacks research entries.
  • [x] GREEN: implement the scan. "Filled-in" needs a definition that survives contact with the template — derive it from divergence from specs/_template/, not from a word count.
  • [x] Distinguish a real marker from a documented one: a placeholder body (…, specific question) is an illustration, not an open question. Without this, any spec that explains the convention reads as permanently blocked.
  • [x] Run against the current repo. Spec 001 must pass both; if it does not, that is a real finding about this spec, not a reason to loosen the test. Passes both.
  • [x] Confirm both have teeth — set the spec back to draft, then plant a real marker; each is named in the failure.
  • [x] Commit.

Verified by: SC7: …, SC8: …, both passing against the live repo.


T9 — Scaffold: pure core ​

Satisfies: R9 (logic half).

  • [x] RED: tests/workflow/scaffold.test.ts — R9: nextSpecNumber allocates the next number, covering empty repo, ["001-a"], gaps, and non-conforming directory names.
  • [x] RED: R9: substitutePlaceholders replaces NNN and slug throughout.
  • [x] GREEN: implement both in scripts/new-spec.ts as pure functions over strings and directory names — no filesystem, no git. Purity is what makes them testable, per plan.md Data & contracts.
  • [x] Confirm the allocation rule has teeth — count from the length instead of the highest, watch it fail, restore.
  • [x] Commit.

Verified by: R9: nextSpecNumber allocates the next number, R9: substitutePlaceholders replaces NNN and slug throughout.


T10 — Scaffold: filesystem, git, and the branch ​

Satisfies: P3 #1, P3 #2, SC4, R9 (shell half).

  • [x] RED: P3 #1: scaffold creates the next numbered directory from the template — against a temporary repo, asserting all four template files land with placeholders substituted.
  • [x] RED: P3 #2: scaffold leaves the working branch at NNN-<slug> — the human confirmed the scaffold creates the branch.
  • [x] GREEN: implement main() — read specs/, call the pure core, copy templates, git checkout -b.
  • [x] Decide and record what happens when the branch already exists. Failing loudly beats silently checking out someone else's work. Refuses, before writing anything, so no half-scaffolded directory is left behind.
  • [x] Copy every .md the template ships rather than a hard-coded list of four, so adding a fifth artifact later cannot be silently skipped.
  • [x] Commit.

Verified by: P3 #1: scaffold creates the next numbered directory from the template, P3 #2: scaffold leaves the working branch at NNN-<slug>, SC4: opening a new feature takes one command.


T11 — /new-spec command wrapper ​

Satisfies: R9 (ergonomics).

  • [x] Add .claude/commands/new-spec.md invoking node scripts/new-spec.ts <slug>.
  • [x] Keep it logic-free. Anything it decides for itself is untestable — that is why T9/T10 exist.
  • [x] Run it end to end in a throwaway repo: four artifacts, placeholders substituted, branch switched.
  • [x] Silence git's expected fatal: from the branch probe, which was leaking to the user's terminal.
  • [x] Commit.

Verified by: covered by T10's tests; the wrapper adds no logic to test. Confirmed by running it once.


T12 — Traceability, and the test that guards it ​

Satisfies: SC5, SC6. Closes the definition of done.

  • [x] Fill the traceability table in spec.md — every acceptance scenario and success criterion mapped to the test that covers it, with proxies marked human-verified and SC1 marked known-unsatisfiable.
  • [x] RED: SC5: every criterion maps to a test that exists — parse the table, assert no empty cells, and assert every test name it references is present in the suite. Self-referential by design: this is what stops the table rotting into decoration.
  • [x] GREEN: reconcile. A criterion with no test is a gap in the work, not a gap in the table. Six real gaps found: P5 #1, #3, #4, #5, #6 had no test at all, and neither did SC6.
  • [x] SC6: no any and no unexplained escape hatches — scan added files, plus pnpm typecheck.
  • [x] Add a Kind column, so a proxy test is never mistaken for proof of behaviour.
  • [x] Full run: pnpm typecheck && pnpm test.
  • [x] Commit.

Verified by: SC5: every criterion maps to a test that exists, SC6: no any and no unexplained escape hatches, and a clean full run.


Notes ​

Staging area for decisions and surprises found during implementation. Move each one into spec.md, research.md, or docs/ before closing the feature — this section is not a home.

  • Bootstrapping. T2 installs the plugin the later tasks are meant to be executed with. T1 and T2 therefore run without the execution skills available; from T3 on, use them. Worth watching whether the difference between the two halves is visible in the diffs — that is the closest thing to a real test of whether this spec was worth building.
  • research.md V7 is open and T2 closes it. It is the only unverified claim left in the plan.
  • Graduation — done. Moved to docs/architecture/workflow-tooling.md at step 6: the two-command setup, the floating version and the pinning routes already tried, the ~800-token SessionStart cost, and Node's unflagged type stripping.