Skip to content

Conversation trace — the natural-language assistant ​

Not a spec. This is the durable record of the brainstorming conversation that led to spec 029, kept so the reasoning survives a context clear. The spec itself is spec.md.

The starting question ​

"I want a prompt in the FE where the user writes 'What is the legendary I am the closest to craft?', and the BE replies — aggregated by an LLM + our BE. How would it work?"

Key realizations ​

  1. The backend already holds the answer. RankingService.rank(apiKey) already returns, per Gen-1 legendary, gatedInputs[] (have/needed/satisfied), myCost (remaining gold once owned mats are subtracted) and craftable. "Closest to craft" is a scoring/sorting question over rows we already produce — not a computation the engine lacks. So the LLM's job is interpret → pick a capability → narrate, never arithmetic.
  2. The MCP server already exposes the tools (spec 017): priory_legendary_ranking, priory_account_materials, priory_recipe_tree, priory_legendaries, with the account key arriving as the X-GW2-Key header, never a tool argument. An external LLM client could already answer this today. The new product work is bringing that loop inside the app.
  3. The LLM is a front door, not a brain. Correctness stays in the deterministic, tested engine — which is exactly the SDD-guardrail sweet spot the project is built around.

The core decision ​

  • Free-text input, not a button. A button ("closest to craft") is a pre-built request parsed by the BE — there is no AI in it. The user writes a sentence; the LLM understands it (synonyms, phrasing we never anticipated, typos, intent) and maps it to a capability we already have. That mapping is the AI, and it is real from the very first stage.
  • The engine stays deterministic either way. "LLM vs deterministic" is therefore not a whole-system switch — it is only a choice about the front door: natural-language input vs structured UI input. The back end is the same tested engine in both.

LLM front door vs deterministic front door — the trade ​

DimensionDeterministic (UI/typed query)LLM (free-text → interpret)
Correctness / testabilityExact, unit/snapshot-testableTest the tools; text→plan can only be eval'd
Cost / latencyZero, instantTokens, round-trips, keys, rate limits
Open-ended questionsCan't — one control per questionThe whole point
Learning value (AI goal)~NoneTool use, agent loop, structured output, reuses 017
Build effort nowLow — engine + endpoints existHigher — endpoint, prompt, guardrails, ops
Failure surface~NoneWrong tool/args, confident-wrong prose, injection

Heuristic kept: if you can imagine the button, build the button; reach for the LLM only when you can't enumerate the buttons. The end vision (ask anything, generate pages) is unbounded → the front door has to be the model.

The iterative staircase (strong base, add complexity when needed) ​

Only the brain behind the seam changes across stages; the seam and the engine never do.

  1. One thing — sentence → LLM understands → maps to one capability (closest_to_craft) → run it → show typed result. Already real AI.
  2. A menu — LLM maps the sentence to any of several capabilities and picks one (intent becomes a discriminated union). Still deterministic control flow.
  3. Combining — for multi-part questions the LLM chains several capabilities itself, deciding the order (the agent loop over the existing MCP tools; add guardrails: allow-list, step cap).
  4. Beyond our data — fuzzy knowledge questions (the wiki — "where's this mastery point?") via GraphRAG, plus richer entity-typed rendering.

What grows across stages is only how far the LLM is allowed to reach: one → menu → combine → combine + world knowledge. Underneath — sentence in, tested engine does the work, typed result out — stays identical. Each stage is its own NNN-<slug> spec; you never need the agent loop to get real AI on screen.

The strong base — "the seam" (what spec 029 is about) ​

Two contracts locked once, extended forever:

text
FE  ──sentence──▶  POST /api/assistant/ask { question: string }
                                   │
                          (the brain: Stage 1 = LLM parses intent)
                                   │
                   deterministic capability layer (e.g. ranking.rank)
                                   │
FE  ◀──typed blocks──  { blocks: AssistantBlock[] }

The return is a developer-owned discriminated union, not free text — a small closed set of render-able blocks the FE knows how to draw:

ts
// design sketch, not repo code
type AssistantBlock =
  | { kind: "prose";            text: string }
  | { kind: "legendaryRanking"; rows: RankingRow[]; highlightId: number }
  | { kind: "recipeTree";       itemId: number }
  | { kind: "gatedGaps";        itemId: number; gaps: GatedInput[] }
  • The LLM selects and annotates; the developer owns the schema and the components; the engine owns the facts. Displayed numbers are the engine's, never the model's.
  • Critical trick: the model returns references (ids), and the server re-hydrates the actual data (e.g. rows) from the tool result it already computed — so the model can never drift a number.
  • Fits the existing pipeline: Zod → nestjs-zod DTO → OpenAPI → Orval client. The same Zod schema (via zod-to-json-schema) becomes the model's structured-output schema — one source of truth for "what shapes can come back."

Locked early (load-bearing, painful to change) ​

  • Endpoint shape: free-text in, typed blocks[] out. The FE binds to this and nothing else.
  • Capability layer as a clean typed boundary — each capability a typed function, logic in services, not controllers. This becomes both the router target and the agent tool later.
  • We own the block/entity union (starts small).

Deferred (swappable per stage, no rewrite) ​

  • Router (stage 1–2) vs agent loop (stage 3); model/provider; @anthropic-ai/sdk vs Vercel AI SDK (note: Vercel's streamUI is RSC/Next-shaped and does not fit our Vite SPA — we can use the pattern and @anthropic-ai/sdk directly, which the brief already commits to); streaming; narration templated vs LLM; number of intents/tools.

Parked for their own focused talks ​

  • Structured LLM replies, in depth (schema design, validation, streaming).
  • FE/BE "build pages structurally" — the generative-UI rendering contract.
  • Design-system + "lego" component approach for entity-typed rendering.
  • Ontology / knowledge graph. GW2 is a strong candidate (rich, stable entity relations). The official GW2 Wiki reportedly runs Semantic MediaWiki [NEEDS VERIFICATION] — part of the knowledge half may already be queryable upstream. Spectrum: bag-of-tools → typed entity/relation layer (the sweet spot; ~80% of the value, no OWL/RDF/reasoner) → full formal ontology (a real learning target, but a whole subsystem). Ontology helps most for the knowledge questions (GraphRAG), least for the computation ones the recipe-graph engine already nails.

Existing assets this builds on ​

  • apps/api/src/legendaries/ranking.service.ts — rank(apiKey) → RankingRow[] (the Stage-1 capability).
  • apps/api/src/legendaries/ranking.schema.ts — RankingRow, GatedInput (Zod, drives OpenAPI/Orval).
  • apps/api/src/recipe-graph/recipe-graph.service.ts — resolvePriced(itemId) (buy-vs-craft tree).
  • apps/api/src/account/account.service.ts — materials/wallet/owned items (needs X-GW2-Key).
  • apps/api/src/mcp/mcp.tools.ts — the tool surface (spec 017) that becomes the agent's tools later.

Open questions to resolve in the spec ​

  • Exact scope of 029: seam + exactly one capability (Stage 1), or a small menu (Stage 2)?
  • What "closest to craft" means precisely — the scoring (gated-inputs-satisfied vs remaining gold vs a blend), and how to rank rows that are already craftable.
  • Narration: templated or a second LLM call, for v1.
  • How the free-text endpoint receives the account key (reuse the X-GW2-Key header convention).
  • How to test a non-deterministic front door (eval the parse; unit-test the capability + block assembly).
  • LLM provider/model + cost/latency guardrails (timeout, max output).