Skip to content

Gaps — the spec workflow ​

Raw observations about how the workflow in CLAUDE.md performed, recorded where they were noticed and deliberately left unprocessed. Nothing here is a proposal, a decision, or a plan. This is a parking lot: evidence goes in, and what to do about it is decided later, separately, with more of it.

One section per feature that produced something worth keeping. Not every feature will.


From spec 010 — frontend conventions (2026-08-05 → 2026-08-08) ​

Nine tasks, 20 commits, superpowers:subagent-driven-development: a fresh implementer per task, a two-stage review after each, one whole-branch review at the end. Roughly 1.5M subagent tokens. Five human decisions were needed mid-implementation, three of them plan conflicts.

Eight defects, and where each was caught ​

Six of the eight were in artifacts that had already passed both the agent's self-review and the human's approval.

#DefectCaught by
1tasks.md's compiler test asserted the plugin's symbols were imported, never that it was wired into the plugins arrayTask review
2tasks.md's ErrorBoundary snippet failed noImplicitOverride: true — it would not compileImplementer
3The SC6 test proved the Zod schema was strict (never in doubt) rather than proving the criterionTask review
4The Biome naming rule would have flagged Orval's generated kebab-case files, which verify:contract forbids touchingController, pre-dispatch
5plan.md's architecture diagram was rejected by the guard rules stated in that same planExecution (T7)
6research.md F6 was false: it claimed Panda emits a semantic token's CSS variable only when referencedExecution (T8)
7Rarity colour values entered through tasks.md unsourced and were never verified in discoveryController fact-check
8The guard suite was narrower than both delivered documents claimedFinal whole-branch review

Defects 1 and 8 needed something to exist and be executed or read adversarially. The other six were present in the artifacts at approval time.

Absence read as proof ​

Defect 6 is the one that generalises. Discovery concluded Panda emits a semantic token's CSS variable only once something references it. The evidence was a grep of a pre-build artifact that found nothing — absence of a string in one file read as proof of conditional behaviour.

It was wrong: every defined token is emitted unconditionally, and the pre-build file simply was not the artifact carrying them. The conclusion it supported happened to survive for a different, correct reason, which is the part worth noticing — a finding that reaches the right answer by the wrong route does not announce itself.

Absence has many explanations: wrong file, wrong timing, wrong command, feature not yet triggered. Only one of them is "the behaviour does not exist."

The two gate types ​

Every gate is a reading of an artifact. Reading caught the structural problems reliably across nine tasks — no missing requirement, no unresolved marker, and zero scope drift, with no implementer building anything the spec had not asked for. It did not catch claims that were simply false about the world: a snippet that would not compile, a value that was not the real value, a rule whose scope was wider than intended.

Spec 011 solved two of the same problems differently ​

Noticed after 010 was complete, on reading main. Spec 011 (apps/api conventions) built:

  • An AST-based guard suite. apps/api/src/conventions/source-model.ts parses with @swc/core and feeds a separate guards.ts with its own tests. 010's guards are regex predicates embedded in a test file. Most of the seven blind spots 010 documented in react.md — dynamic import(), path aliases, re-export chains, Component:/lazy: route forms — are artifacts of matching text rather than parsing it.
  • A doc test. tests/docs/nestjs-guide.test.ts asserts the guide exists and is linked from CLAUDE.md. Minimal — it does not check the document's claims — but 010's equivalent criterion (SC8) is verified by human review only.

011's own tasks.md Notes record two findings its gates missed and implementation found: import.meta rejected by tsc in apps/api (TS1470), and a pnpm install prerequisite before any api test runs.

010's agent did not read 011's approach before designing 010's guards, though 011 was merged to main while 010 was in flight. That is an observation about the agent, not the workflow.

Cost ​

Subagent-driven development ran roughly 1.5M subagent tokens; each fix round costs a dispatch plus a scoped re-review. What it bought that inline execution could not: implementers with no memory of authoring the artifacts they were checking against, and a final reviewer reading nineteen commits as one diff. Defects 2, 5 and 8 are directly attributable to that separation.

Left unprocessed ​

Listed so they are not lost, not because any of them is recommended:

  • Code blocks in tasks.md are approved without being typechecked (defect 2).
  • Factual values can enter a spec or task list without tracing to research.md (defect 7).
  • No gate checks a plan against rules the plan itself states (defect 5).
  • A negative discovery finding is not required to say what a positive would have looked like (defect 6).
  • Tests in tasks.md are not routinely asked what would make them fail (defects 1 and 3).
  • Two convention-enforcement implementations now exist in the repo, of differing approach and quality (010 vs 011).