Natural-language assistant (seam + closest-to-craft) — Implementation Plan
Status: approved Status is set by the human, never by the agent. plan.md must be approved before tasks.md (Step 3). Approved by the human on 2026-08-20, transcribed by the agent on their explicit "go" (parse model claude-haiku-4-5 accepted). The agent did not set this on its own initiative.
Scope of this file (constitution override of
superpowers:writing-plans). This is the header half only: goal, architecture, global constraints, file structure, and a task outline with locked interfaces. The bite-sized TDD steps (test code, commands, commits) live intasks.md, written in Step 3 after this plan is approved. No execution handoff here — implementation is entered only through plan mode (Step 4).
Goal: Ship the assistant seam — a free-text question in, a typed answer out — proven by one live capability, "closest to craft", rendered abstractly on the web.
Architecture: A new assistant NestJS feature module exposes POST /api/assistant/ask. Its service makes exactly one LLM call — through an injectable, mockable @anthropic-ai/sdk wrapper — to classify the free-text question into a Zod discriminated intent union (closest_to_craft | unsupported). A closest_to_craft intent runs the deterministic capability over the existing RankingService, sorts by remaining gold, and the service returns a typed intent-discriminated answer envelope that reuses RankingRow. A minimal web feature renders the envelope abstractly (one-line summary + a plain list, or an honest decline), authenticating with the existing Authorization: Bearer key flow. The LLM interprets and narrates only; every displayed number comes from tested code.
Tech Stack: NestJS 11 + Fastify, nestjs-zod + zod (both already in the graph), @anthropic-ai/sdk (the only new dependency — research F4), the Zod→OpenAPI→Orval client pipeline, React (Vite SPA) + TanStack Query, Vitest + Testing Library, Biome.
Global Constraints
Project-wide requirements; every task's requirements implicitly include these. Values copied verbatim.
- One LLM call per request. The intent parse is the only model call; the
summaryis templated (spec R3/R5, SC6). - The LLM never produces displayed facts. All
RankingRownumbers come fromRankingService; the model returns an intent only (spec R4/R6). - Auth =
Authorization: Bearer, reusinglegendaries.controller.tsrequireBearer+mapGw2Errorexactly:400missing/blank,401invalid,403missing scope,502upstream; the key is never logged or echoed (spec R2, C2, SC4). - Zod-first contract, reusing the existing
RankingRowschema (no parallel copy); the response DTO drives@ZodResponse→ OpenAPI → Orval (spec R6, SC5). - Structured output via
client.messages.parse({ output_config: { format: zodOutputFormat(AssistantIntent) } }); guardparsed_outputfor null (research V1/F4). - Model id (research F5, human decision to confirm in Task 4): default
claude-haiku-4-5— cheap, low-latency, supports structured output; escalate toclaude-opus-4-8only if the eval shows the parse is weak. Notclaude-opus-4-8-by-reflex. - Config via
process.envthrough a pure resolver mirroringconfig/port.ts; fail fast at boot whenANTHROPIC_API_KEYis absent (research V2). - The LLM client is injectable and mocked in controller/service tests — no live model, no tokens in CI (spec R9).
- NLU quality is an eval, not a CI gate (spec C1):
intent.eval.tsruns against the real model on demand, never inpnpm test. - TypeScript: no
any, no unexplained escape hatches;exactOptionalPropertyTypesis on — optional props are omitted, neverundefined; Nest DI needs value imports for injected services (docs/architecture/nestjs.md; constitution DoD). - Web design system: no new tokens; gold renders through the spec-020 coin component, never a literal; the tokens-never-literals guard scans comments too — no
rgb(/hsl(/#hexanywhere (docs/architecture/design-system.md; memory). - React Compiler build gate: avoid destructured-default + inline-type props;
pnpm build(not vitest) catches it (memory). - Green before push:
pnpm lint(Biome),pnpm test,pnpm build,pnpm docs:buildall pass (spec SC8). Node is at/opt/homebrew/bin(prefix PATH); pnpm 11.15.1. - Nothing written to
docs/superpowers/— a test asserts the count stays zero (spec R10).
File Structure
Files that change together live together in apps/api/src/assistant/ and apps/web/src/features/assistant/.
API — apps/api/src/
config/anthropic-key.ts— Create.resolveAnthropicKey(env: NodeJS.ProcessEnv): string, mirrorsport.ts; throws a clear error when unset (fail-fast at boot).assistant/assistant.schema.ts— Create. Zod source of truth:AskRequest,AssistantIntent,AssistantAnswer(reusingRankingRowfrom../legendaries/ranking.schema), pluscreateZodDtoexports. One responsibility: the contract.assistant/closest-to-craft.ts— Create. Pure: the comparator, the sort, the templated summary. No I/O — golden-value tested.assistant/anthropic.client.ts— Create. Injectable wrapper over@anthropic-ai/sdk; the only place the SDK is imported. Constructed fromresolveAnthropicKey(fail-fast). ExposesparseIntent. Mockable by DI class token.assistant/assistant.service.ts— Create. Orchestrator: parse → route → envelope. Depends onAnthropicClient+RankingService.assistant/assistant.controller.ts— Create. Thin:POST /assistant/ask,@Headers('authorization'),requireBearer, error mapping,@ZodResponse. No logic.assistant/assistant.module.ts— Create. Providers (AnthropicClient,AssistantService) + import the module that providesRankingService; export nothing.app.module.ts— Modify. RegisterAssistantModule.assistant/eval/intent.eval.ts— Create. Sentences → expected intent, run against the real model on demand. Excluded from the CI test glob (C1).package.json(apps/api) — Modify. Add@anthropic-ai/sdk.- Tests (co-located):
assistant.schema.test.ts,closest-to-craft.test.ts,assistant.service.test.ts(mocked client),assistant.controller.test.ts,assistant.module.test.ts.
Web — apps/web/src/
- Generated Orval client — Regenerate (not hand-edited) so the
askmutation +AssistantAnswertypes exist. features/assistant/AssistantView.tsx— Create. Presentational: free-text input + submit; renderssummary+ a plain results list (name + gold via the coin component), or theunsupportedmessage.features/assistant/AssistantPage.tsx— Create. Wires the mutation +useApiKey; no key →ConnectAccountPrompt(reuse spec 016/020).features/assistant/routes.tsx— Create./assistant→AssistantPage(feature owns its route table —docs/architecture/react.md).- App shell nav — Modify. One link to
/assistant. - Tests:
features/assistant/__tests__/AssistantView.test.tsx,AssistantPage.test.tsx.
Locked interfaces
Every task implements against these exact names/types; later tasks rely on them.
// assistant.schema.ts
type AskRequest = { question: string }
type AssistantIntent =
| { intent: 'closest_to_craft' }
| { intent: 'unsupported'; reason: string }
type AssistantAnswer =
| { intent: 'closest_to_craft'; summary: string; results: RankingRow[] }
| { intent: 'unsupported'; message: string }
// closest-to-craft.ts (pure)
// myCost asc, null last, id asc tiebreak — mirrors ranking.service.ts byPersonalProfitDescIdAsc
function byRemainingGoldAscIdAsc(a: Pick<RankingRow,'id'|'myCost'>, b: Pick<RankingRow,'id'|'myCost'>): number
function closestToCraft(rows: RankingRow[]): RankingRow[] // returns a sorted copy
function buildSummary(rows: RankingRow[]): string // top row, or "nothing rankable yet"
// anthropic.client.ts (injectable; the only @anthropic-ai/sdk import)
class AnthropicClient { parseIntent(question: string): Promise<AssistantIntent> }
// assistant.service.ts
class AssistantService { ask(question: string, apiKey: string): Promise<AssistantAnswer> }Task outline
Ordered; each ends with an independently testable deliverable and its own commit. Detailed TDD steps are written into tasks.md at Step 3.
- T1 — Dependency + config resolver. Add
@anthropic-ai/sdk; createconfig/anthropic-key.ts(resolveAnthropicKey, fail-fast). (spec R8; research V2/F4) - T2 — Contract schemas.
assistant.schema.ts:AskRequest,AssistantIntent,AssistantAnswer(reusingRankingRow), DTOs; schema test asserts the discriminated union validates and rejects a bad shape. (spec R6; SC5) - T3 — Deterministic capability.
closest-to-craft.ts: comparator +closestToCraft+buildSummary, golden-value tests (myCost asc, null last, id tiebreak; empty/all-null → nothing-rankable). (spec R4/R5; SC3) - T4 — Anthropic client wrapper.
anthropic.client.ts:parseIntentviamessages.parse+zodOutputFormat(AssistantIntent); pin the model id (F5 decision); one throwaway real call in dev to confirmparsed_outputvalidates, then rely on the mock. (spec R3/R8; research V1/F4/F5) - T5 — Service orchestration.
assistant.service.ts: parse →closest_to_craft(rank → sort → summary → envelope) |unsupported(decline). Tests with the client mocked: ordering, one-call assertion, unsupported returns no results and does not callRankingService(spy). (spec R3–R6; P1 #1/#3/#4, P2 #1/#3; SC1/SC2/SC6) - T6 — Controller + module wiring.
assistant.controller.ts(POST /assistant/ask, Bearer, error map),assistant.module.ts, register inapp.module.ts; controller test covers the400/401/403/502mapping and key-never-echoed; module/arch tests stay green. (spec R1/R2, C2; SC4) - T7 — OpenAPI + Orval regen. Regenerate the OpenAPI document and the web client so the
askmutation +AssistantAnswertypes exist; assert the generated client is present. (spec R6; SC5) - T8 — Minimal web feature.
AssistantView+AssistantPage+ route + nav link; abstract render (summary + list via coin component, or message); no-key →ConnectAccountPrompt; view/page tests with the mutation hook stubbed (ranked / declined / no-key). No new tokens; guards green. (spec R7; P1 #2, SC7) - T9 — Eval + verification.
intent.eval.ts(excluded from CI); fill the traceability table; runpnpm lint/test/build/docs:buildgreen. (spec R9/R10, C1; SC8)
Self-review (against the spec)
- Coverage: every spec R and SC maps to a task above — R1/R2→T6, R3→T4/T5, R4/R5→T3/T5, R6→T2/T7, R7→T8, R8→T1/T4, R9→T5/T9, R10→T9; SC1/2/6→T5, SC3→T3, SC4→T6, SC5→T2/T7, SC7→T8, SC8→T9. No orphan requirement.
- Placeholders: none — the detailed code is deferred to
tasks.mdby design (constitution split), not left as "TODO". - Type consistency:
AssistantIntent/AssistantAnswer/parseIntent/closestToCraft/buildSummary/askare used with identical names and signatures across T2–T8 (see Locked interfaces).