Skip to content

Spec 029 — Natural-language assistant (the seam + closest-to-craft) ​

Status: implemented Branch: 029-nl-assistant

Status is set by the human, never by the agent. It moves draft → approved → implemented. Approved by the human on 2026-08-20, transcribed by the agent on their explicit "go" — after research.md closed V1–V3 and C1/C2 were resolved. The agent did not set this on its own initiative. Implemented on the human's decision (2026-08-20), transcribed into the branch before merge so the status transition is reviewed and merged with the work (Definition of Done). All nine tasks landed with per-task tests; the live NLU eval passed 18/18 and the endpoint was smoke-tested end-to-end against real Haiku 4.5.

Problem ​

The tool answers questions only through fixed UI: you navigate to a page, or you hit a typed endpoint. There is no way to ask in plain language. The project brief names an LLM layer that "parses a fuzzy goal into structured constraints" as the intended agentic surface, and spec 017 already exposes the deterministic engine (ranking, recipe trees, account reads) as MCP tools — but nothing inside the product lets a player type a sentence and get an answer. A button labelled "closest to craft" would be a pre-built request parsed by the backend; there is no language understanding in it. We want the player to write a sentence, have an LLM understand it, and map it to a capability the engine already computes — with every displayed number coming from that tested engine, never from the model.

Approach — the seam (why this shape) ​

This spec builds the strong base every later assistant stage stands on: a stable contract of sentence in → typed result out, with the LLM as a thin front door (it interprets and routes; it never computes). Only the "brain" behind this seam grows in later specs (a menu of capabilities, then a tool-use loop, then world-knowledge). The full design rationale — the LLM-vs-deterministic trade, the four-stage staircase, the block-union return contract, and the parked topics (generative UI, ontology / GraphRAG) — is recorded in conversation-trace.md alongside this file. This spec deliberately implements Stage 1 only: one live capability behind the seam, rendered abstractly.

HTTP endpoints (contract-first) ​

One new endpoint. Its request and response are Zod DTOs that drive @ZodResponse, the OpenAPI document, and therefore Orval's generated web client — the same pipeline every other endpoint uses (spec 023 era).

  • POST /api/assistant/ask — the single assistant entry point.
    • Auth: Authorization: Bearer <key> — the same convention GET /api/legendaries/ranking uses (legendaries.controller.ts requireBearer). The key is required up front — C2, resolved — because the only capability (closest_to_craft) is inherently personal: it reads the account's owned materials. The key is never logged or echoed.

    • Request body: { question: string } — non-empty, length-bounded.

    • Response 200: a discriminated union on intent (R6):

      ts
      // assistant.schema.ts — the response contract (design sketch)
      { intent: "closest_to_craft"; summary: string; results: RankingRow[] }
      | { intent: "unsupported";      message: string }

      RankingRow is the existing schema (ranking.schema.ts), reused — not re-declared.

    • Errors: reuse the ranking endpoint's mapping exactly — missing/blank key → 400, invalid / expired key → 401, key missing a scope → 403, GW2 upstream failure → 502. An LLM/provider failure maps to 502 (or a dedicated 5xx), never leaking the key or prompt internals.

No other endpoint or response shape changes.

User stories ​

Ordered by priority. Each is independently testable and shippable — if only P1 ships, there is still a usable, valuable feature.

P1 — Ask "which legendary am I closest to crafting?" in plain language ​

As a connected player, I want to type a question in my own words — "which shiny am I nearly done with?", "what legendary do I have the most mats for?" — and get my Gen-1 legendaries ranked by how little gold I have left to spend, so that I learn what to finish next without navigating menus.

Independent test: with a stored key, submitting a free-text question that means "closest to craft" (several distinct phrasings) to POST /api/assistant/ask returns intent: "closest_to_craft" with a one-line summary and results = the account's Gen-1 legendaries ordered by myCost ascending (null last); the web assistant box renders that summary + a plain list (name + remaining gold). Every number in results comes from RankingService, not the model.

Acceptance scenarios

  1. Given a stored key and a question meaning "closest to craft" (any phrasing), whenPOST /api/assistant/ask runs, then the LLM classifies it closest_to_craft and the response carries results ordered by myCost ascending with null-myCost rows last, ties broken by id.
  2. Given that response, when the web assistant box renders it, then it shows the summary line and a plain list of results (each: legendary name + remaining gold via the coin component) — no block-kind components.
  3. Given a closest_to_craft result, when the service builds the summary, then the summary is composed deterministically from the top row (no second LLM call); an empty / all-null list yields a "nothing rankable yet" summary rather than a crash or a fabricated pick.
  4. Given exactly one request, when it is served, then exactly one LLM call is made (the intent parse); the ranking and the summary use no further model calls.

P2 — Off-topic questions are declined honestly (no hallucination) ​

As a player, I want the assistant to admit when it cannot answer — "where's this mastery point?", "is Bifrost profitable?" — rather than inventing an answer or dumping an unrelated ranking, so that I trust what it does say.

Independent test: submitting a question the single capability cannot serve returns intent: "unsupported" with a short, honest message that names what the assistant can do; no ranking is returned and no world-knowledge answer is fabricated. The web box renders that message.

Acceptance scenarios

  1. Given an off-topic question, when POST /api/assistant/ask runs, then the LLM classifies it unsupported and the response is { intent: "unsupported", message } — no results, no invented answer.
  2. Given an unsupported classification, when the response is built, then message is honest and actionable (states the assistant currently only answers "closest to craft"), derived from the model's reason or a templated default.
  3. Given an unsupported question, when it is served, then the ranking capability and RankingService are not invoked (no account read for a question that does not need one).

Requirements ​

  • R1 — New assistant feature module. Add apps/api/src/assistant/ per nestjs.md (feature-module layout, thin controller, Zod-first contract, service owns the work): assistant.controller.ts, assistant.service.ts, assistant.schema.ts, assistant.module.ts, wired into AppModule. The controller only reads the header + body and delegates; orchestration lives in the service.
  • R2 — Endpoint POST /api/assistant/ask. Body { question: string } (non-empty, length-bounded via the request DTO; over-long or empty → 400 before any model call). Key via Authorization: Bearer, parsed with the ranking endpoint's requireBearer semantics, key never logged. Error mapping identical to the ranking route (400/401/403/502). Returns the R6 union on success.
  • R3 — Intent parse (the front door, one model call). AssistantService sends the question to an LLM via the R8 client and receives a value conforming to the Zod discriminated union AssistantIntent = { intent: "closest_to_craft" } | { intent: "unsupported"; reason: string }, using the SDK's structured-output mechanism — messages.parse + output_config.format with the SDK's zodOutputFormat over that one Zod definition (confirmed, research V1/F4; the helper removes the need for a separate schema-conversion library). Exactly one model call per request. The model is instructed: choose closest_to_craft for any request to identify which legendary the account is nearest to crafting, in any phrasing; otherwise unsupported with a short user-facing reason. No parameters are parsed in this spec (pure classification). The union is the extension seam — later specs add variants, additively.
  • R4 — Capability closest_to_craft (deterministic). A service method calls RankingService.rank(key) → RankingRow[], then orders by a named comparator = myCost ascending, rows with myCost === null last, ties broken by id ascending (mirrors the null-last convention in ranking.service.tsbyPersonalProfitDescIdAsc). Returns the ordered rows in the existing RankingRow shape, closest first. The comparator is isolated so a future time/annoyance term extends the sort without touching the endpoint or the response union. Design note (future, not built here): that later term must treat gated inputs as fungible (an owned Gift of Exploration counts once and applies across all legendaries) — modelled as "distinct gated acquisitions still needed", not per-legendary gated counts.
  • R5 — Templated summary (no second model call). For a closest_to_craft result the service composes the one-line summary deterministically from the top row (legendary name + remaining gold), or a "nothing rankable yet" line when the list is empty or every myCost is null. No LLM call is made to write it.
  • R6 — Response contract (Zod union → OpenAPI → Orval). assistant.schema.ts defines an intent-discriminated Zod union, exported as a createZodDto, reusing the existing RankingRow schema (no parallel copy): { intent: "closest_to_craft"; summary; results: RankingRow[] } or { intent: "unsupported"; message }. It drives @ZodResponse, the OpenAPI document, and Orval's generated web client. message carries the honest decline text (model reason or a templated default).
  • R7 — Minimal web feature (free-text box). Add apps/web/src/features/assistant/ per react.md (feature-folder layout, Page/View naming): a single free-text input + submit that calls POST /api/assistant/ask through the generated Orval mutation hook (a POST/mutation, not a Suspense query — pending/error handled locally, not via QueryBoundary). It renders the response abstractly: closest_to_craft → the summary + a plain list of results (name + remaining gold via the spec-020 coin component); unsupported → the message. The Bearer key comes from the existing useApiKey(); with no stored key the box shows the shared ConnectAccountPrompt (the capability is account-dependent), reusing specs 016/020. No new design tokens.
  • R8 — LLM client boundary + guardrails. The @anthropic-ai/sdk call sits behind an injectable client wrapper (so tests mock it and later stages swap it), reading its API key from process.env via a resolveAnthropicKey resolver mirroring config/port.ts (confirmed, research V2; @anthropic-ai/sdk is the only new dependency, research F4). Guardrails: a request timeout and a bounded max output; a provider/model failure surfaces as a 502-class error that leaks neither the key nor the raw prompt. The model id (a small, fast Claude) is fixed in plan.md.
  • R9 — Testing seams. The R8 client wrapper is injectable and mocked in controller/service tests, so the deterministic pipeline (routing → capability → envelope) is tested with no live model and no tokens. The comparator + summary + capability are pure and unit-tested with golden values. NLU quality (phrasing → intent) is an eval artifact (example sentences → expected intent) run against the real model — eval-only, not a CI gate (C1, resolved).
  • R10 — Guards stay green. No new docs/superpowers/ artifacts; the tokens-never-literals guard, the docs/superpowers/ count guard, the api/web conventions tests, and lint stay green; pnpm docs:build compiles this spec and conversation-trace.md.

Mark anything unresolved inline rather than assuming an answer. Two markers, split by who can answer:

  • [NEEDS CLARIFICATION: specific question] — only the human can answer. Blocks step 1.5.
  • [NEEDS VERIFICATION: specific question] — only reality can answer, resolved in research.md with cited evidence. Blocks the approval gate.

[NEEDS CLARIFICATION] — resolved by the human (2026-08-20), transcribed on their "go":

  • C1 — NLU gating → eval-only. Parse quality is held by a manually / periodically-run eval set (example sentences → expected intent), not a CI gate — no per-run token cost or flakiness in CI. The deterministic contract is still unit-tested with the model mocked.
  • C2 — Key-required posture → require up front. POST /api/assistant/ask requires Authorization and returns 400 when it is absent (matches the ranking route; the sole capability is account-only). Deferring the key to parse-time is a possible later refinement, not MVP.

[NEEDS VERIFICATION] — all resolved before approval; verdicts with cited evidence in research.md (2026-08-20). No marker remains open:

  • V1 — Confirmed (research V1/F4). @anthropic-ai/sdk is the only new dependency; structured output uses messages.parse + zodOutputFormat over the one Zod schema — no separate schema lib. (R3)
  • V2 — Confirmed (research V2). Config is read from process.env via a resolveAnthropicKey resolver like config/port.ts; no @nestjs/config. (R8)
  • V3 — Confirmed (research V3). RankingService.rank() inherits the ~31 s floor (19-character account, spec 017), cached after the first read; the single LLM call is cheap by comparison. A "still working" / timeout note is a plan concern. (R4/R8)

Success criteria ​

Measurable and technology-agnostic — outcomes, not implementation. Deterministic behaviours are asserted with the LLM client mocked; NLU phrasing robustness is an eval (C1).

  • SC1 — A free-text question meaning "closest to craft" is classified closest_to_craft, and the response's results are the account's Gen-1 legendaries ordered by myCost ascending (null last, id tiebreak). (service/controller unit test with the client mocked to return the intent; eval for phrasing)
  • SC2 — An off-topic question is classified unsupported and returns { intent, message } with an honest, actionable message and no results; RankingService is not invoked. (unit test with mocked intent + a spy asserting no ranking call; eval)
  • SC3 — closest_to_craft orders RankingRows by myCost ascending with null last and id tiebreak, and returns the existing RankingRow shape unchanged. (golden-value unit test)
  • SC4 — The endpoint enforces Authorization: Bearer with the ranking route's mapping (400 missing/blank, 401 invalid, 403 missing scope, 502 upstream); the key never appears in logs or the response body. (controller test)
  • SC5 — The response validates as the intent-discriminated union against the DTO / OpenAPI, reusing RankingRow; Orval generates a typed client for it. (schema test + generated-client presence)
  • SC6 — Exactly one LLM call is made per request and the summary is produced with no second model call. (unit test asserting the mocked client was called once)
  • SC7 — In the running app: a plain-language "closest to craft" question shows the ranked answer; an off-topic question shows the decline message; with no stored key the box shows ConnectAccountPrompt. (web tests with the mutation hook stubbed + a step-5 manual run)
  • SC8 — Typecheck, pnpm test, pnpm lint, pnpm build, and pnpm docs:build are green, and every acceptance scenario and success criterion maps to a named test in the traceability table with no gap.

Out of scope ​

  • The developer-owned block-kind union + per-kind render components — the response ships as a generic envelope rendered as a plain list; real block components are the next spec.
  • A multi-intent menu / a second capability — one intent (closest_to_craft) plus unsupported.
  • The agent tool-use loop (Stage 3) and wiki / world-knowledge questions (Stage 4 / GraphRAG).
  • The time/annoyance closeness metric and the fungible-gated model — a future comparator extension; the design intent is recorded in R4, not built here.
  • Intent parameters (limit, generation, weapon filter) — pure classification for MVP.
  • Streaming / SSE responses — a single JSON response.
  • Conversation history / multi-turn memory — each request is stateless.
  • Rich rendering of the ranking (tables, tree, highlighting) — abstract list only.
  • A second LLM call for narration — the summary is templated (R5).

Assumptions ​

  • The ranking engine stands — RankingService.rank(key) returns RankingRow[] with myCost and the null-last convention (specs 018/020/023). Consumed unchanged; no ranking logic is added here.
  • The Bearer auth flow stands — the ranking endpoint's Authorization: Bearer convention (requireBearer, mapGw2Error) and the web useApiKey / ConnectAccountPrompt (specs 015/016/020).
  • The Zod → nestjs-zod DTO → OpenAPI → Orval pipeline stands — the response DTO drives the generated client; reusing RankingRow avoids a parallel schema (spec 023 era).
  • nestjs.md / react.md / design-system.md hold — feature-module + feature-folder layout, thin controllers, Zod-first contracts, Suspense-vs-mutation data flow, tokens-never-literals.
  • The brief's LLM posture holds — @anthropic-ai/sdk, latest small/fast Claude, the LLM at the edges over a deterministic engine; correctness stays in tested code.

Traceability ​

Each acceptance scenario and success criterion maps to a named test; SC8 asserts no empty cell. Filled in during implementation. api/* paths under apps/api/src/, web/* under apps/web/src/.

CriterionTest
P1 #1api/assistant/assistant.service.test.ts — "P1 #1/SC1: closest_to_craft orders results by myCost asc, null last"
P1 #2web/features/assistant/__tests__/AssistantView.test.tsx — "P1 #2/SC7: renders the summary + a plain results list"
P1 #3api/assistant/assistant.service.test.ts — "P1 #3/SC6: summary is templated from the top row…" + "P1 #3: empty ranking → nothing-rankable summary"
P1 #4api/assistant/assistant.service.test.ts — "P1 #3/SC6: …and exactly ONE model call is made"
P2 #1api/assistant/assistant.service.test.ts — "P2 #1/SC2: unsupported returns a message and no results"
P2 #2api/assistant/assistant.service.test.ts — "P2 #1/SC2: unsupported returns a message…" (asserts the message content)
P2 #3api/assistant/assistant.service.test.ts — "P2 #3/SC2: unsupported does NOT invoke RankingService" (spy)
SC1api/assistant/assistant.service.test.ts (mocked intent) + api/assistant/intent.eval.ts (phrasing eval)
SC2api/assistant/assistant.service.test.ts (mocked intent + no-ranking spy) + api/assistant/intent.eval.ts
SC3api/assistant/closest-to-craft.test.ts — golden-value ordering (myCost asc, null last, id tiebreak)
SC4api/assistant/assistant.controller.test.ts — Bearer 400/401/403/502 mapping; key never echoed
SC5api/assistant/assistant.schema.test.ts — union validates; web/api/generated/endpoints/assistant/* present
SC6api/assistant/assistant.service.test.ts — "P1 #3/SC6: …exactly ONE model call is made"
SC7web/features/assistant/__tests__/AssistantView.test.tsx + AssistantPage.test.tsx (ranked / declined / no-key) + step-5 manual run
SC8this table complete + pnpm typecheck/lint/test/build/docs:build/verify:contract green