Gaps — the spec workflow
Raw observations about how the workflow in CLAUDE.md performed, recorded where they were noticed and deliberately left unprocessed. Nothing here is a proposal, a decision, or a plan. This is a parking lot: evidence goes in, and what to do about it is decided later, separately, with more of it.
One section per feature that produced something worth keeping. Not every feature will.
From spec 010 — frontend conventions (2026-08-05 → 2026-08-08)
Nine tasks, 20 commits, superpowers:subagent-driven-development: a fresh implementer per task, a two-stage review after each, one whole-branch review at the end. Roughly 1.5M subagent tokens. Five human decisions were needed mid-implementation, three of them plan conflicts.
Eight defects, and where each was caught
Six of the eight were in artifacts that had already passed both the agent's self-review and the human's approval.
| # | Defect | Caught by |
|---|---|---|
| 1 | tasks.md's compiler test asserted the plugin's symbols were imported, never that it was wired into the plugins array | Task review |
| 2 | tasks.md's ErrorBoundary snippet failed noImplicitOverride: true — it would not compile | Implementer |
| 3 | The SC6 test proved the Zod schema was strict (never in doubt) rather than proving the criterion | Task review |
| 4 | The Biome naming rule would have flagged Orval's generated kebab-case files, which verify:contract forbids touching | Controller, pre-dispatch |
| 5 | plan.md's architecture diagram was rejected by the guard rules stated in that same plan | Execution (T7) |
| 6 | research.md F6 was false: it claimed Panda emits a semantic token's CSS variable only when referenced | Execution (T8) |
| 7 | Rarity colour values entered through tasks.md unsourced and were never verified in discovery | Controller fact-check |
| 8 | The guard suite was narrower than both delivered documents claimed | Final whole-branch review |
Defects 1 and 8 needed something to exist and be executed or read adversarially. The other six were present in the artifacts at approval time.
Absence read as proof
Defect 6 is the one that generalises. Discovery concluded Panda emits a semantic token's CSS variable only once something references it. The evidence was a grep of a pre-build artifact that found nothing — absence of a string in one file read as proof of conditional behaviour.
It was wrong: every defined token is emitted unconditionally, and the pre-build file simply was not the artifact carrying them. The conclusion it supported happened to survive for a different, correct reason, which is the part worth noticing — a finding that reaches the right answer by the wrong route does not announce itself.
Absence has many explanations: wrong file, wrong timing, wrong command, feature not yet triggered. Only one of them is "the behaviour does not exist."
The two gate types
Every gate is a reading of an artifact. Reading caught the structural problems reliably across nine tasks — no missing requirement, no unresolved marker, and zero scope drift, with no implementer building anything the spec had not asked for. It did not catch claims that were simply false about the world: a snippet that would not compile, a value that was not the real value, a rule whose scope was wider than intended.
Spec 011 solved two of the same problems differently
Noticed after 010 was complete, on reading main. Spec 011 (apps/api conventions) built:
- An AST-based guard suite.
apps/api/src/conventions/source-model.tsparses with@swc/coreand feeds a separateguards.tswith its own tests. 010's guards are regex predicates embedded in a test file. Most of the seven blind spots 010 documented inreact.md— dynamicimport(), path aliases, re-export chains,Component:/lazy:route forms — are artifacts of matching text rather than parsing it. - A doc test.
tests/docs/nestjs-guide.test.tsasserts the guide exists and is linked fromCLAUDE.md. Minimal — it does not check the document's claims — but 010's equivalent criterion (SC8) is verified by human review only.
011's own tasks.md Notes record two findings its gates missed and implementation found: import.meta rejected by tsc in apps/api (TS1470), and a pnpm install prerequisite before any api test runs.
010's agent did not read 011's approach before designing 010's guards, though 011 was merged to main while 010 was in flight. That is an observation about the agent, not the workflow.
Cost
Subagent-driven development ran roughly 1.5M subagent tokens; each fix round costs a dispatch plus a scoped re-review. What it bought that inline execution could not: implementers with no memory of authoring the artifacts they were checking against, and a final reviewer reading nineteen commits as one diff. Defects 2, 5 and 8 are directly attributable to that separation.
Left unprocessed
Listed so they are not lost, not because any of them is recommended:
- Code blocks in
tasks.mdare approved without being typechecked (defect 2). - Factual values can enter a spec or task list without tracing to
research.md(defect 7). - No gate checks a plan against rules the plan itself states (defect 5).
- A negative discovery finding is not required to say what a positive would have looked like (defect 6).
- Tests in
tasks.mdare not routinely asked what would make them fail (defects 1 and 3). - Two convention-enforcement implementations now exist in the repo, of differing approach and quality (010 vs 011).