Conversation trace — the natural-language assistant
Not a spec. This is the durable record of the brainstorming conversation that led to spec 029, kept so the reasoning survives a context clear. The spec itself is
spec.md.
The starting question
"I want a prompt in the FE where the user writes 'What is the legendary I am the closest to craft?', and the BE replies — aggregated by an LLM + our BE. How would it work?"
Key realizations
- The backend already holds the answer.
RankingService.rank(apiKey)already returns, per Gen-1 legendary,gatedInputs[](have/needed/satisfied),myCost(remaining gold once owned mats are subtracted) andcraftable. "Closest to craft" is a scoring/sorting question over rows we already produce — not a computation the engine lacks. So the LLM's job is interpret → pick a capability → narrate, never arithmetic. - The MCP server already exposes the tools (spec 017):
priory_legendary_ranking,priory_account_materials,priory_recipe_tree,priory_legendaries, with the account key arriving as theX-GW2-Keyheader, never a tool argument. An external LLM client could already answer this today. The new product work is bringing that loop inside the app. - The LLM is a front door, not a brain. Correctness stays in the deterministic, tested engine — which is exactly the SDD-guardrail sweet spot the project is built around.
The core decision
- Free-text input, not a button. A button ("closest to craft") is a pre-built request parsed by the BE — there is no AI in it. The user writes a sentence; the LLM understands it (synonyms, phrasing we never anticipated, typos, intent) and maps it to a capability we already have. That mapping is the AI, and it is real from the very first stage.
- The engine stays deterministic either way. "LLM vs deterministic" is therefore not a whole-system switch — it is only a choice about the front door: natural-language input vs structured UI input. The back end is the same tested engine in both.
LLM front door vs deterministic front door — the trade
| Dimension | Deterministic (UI/typed query) | LLM (free-text → interpret) |
|---|---|---|
| Correctness / testability | Exact, unit/snapshot-testable | Test the tools; text→plan can only be eval'd |
| Cost / latency | Zero, instant | Tokens, round-trips, keys, rate limits |
| Open-ended questions | Can't — one control per question | The whole point |
| Learning value (AI goal) | ~None | Tool use, agent loop, structured output, reuses 017 |
| Build effort now | Low — engine + endpoints exist | Higher — endpoint, prompt, guardrails, ops |
| Failure surface | ~None | Wrong tool/args, confident-wrong prose, injection |
Heuristic kept: if you can imagine the button, build the button; reach for the LLM only when you can't enumerate the buttons. The end vision (ask anything, generate pages) is unbounded → the front door has to be the model.
The iterative staircase (strong base, add complexity when needed)
Only the brain behind the seam changes across stages; the seam and the engine never do.
- One thing — sentence → LLM understands → maps to one capability (
closest_to_craft) → run it → show typed result. Already real AI. - A menu — LLM maps the sentence to any of several capabilities and picks one (intent becomes a discriminated union). Still deterministic control flow.
- Combining — for multi-part questions the LLM chains several capabilities itself, deciding the order (the agent loop over the existing MCP tools; add guardrails: allow-list, step cap).
- Beyond our data — fuzzy knowledge questions (the wiki — "where's this mastery point?") via GraphRAG, plus richer entity-typed rendering.
What grows across stages is only how far the LLM is allowed to reach: one → menu → combine → combine + world knowledge. Underneath — sentence in, tested engine does the work, typed result out — stays identical. Each stage is its own NNN-<slug> spec; you never need the agent loop to get real AI on screen.
The strong base — "the seam" (what spec 029 is about)
Two contracts locked once, extended forever:
FE ──sentence──▶ POST /api/assistant/ask { question: string }
│
(the brain: Stage 1 = LLM parses intent)
│
deterministic capability layer (e.g. ranking.rank)
│
FE ◀──typed blocks── { blocks: AssistantBlock[] }The return is a developer-owned discriminated union, not free text — a small closed set of render-able blocks the FE knows how to draw:
// design sketch, not repo code
type AssistantBlock =
| { kind: "prose"; text: string }
| { kind: "legendaryRanking"; rows: RankingRow[]; highlightId: number }
| { kind: "recipeTree"; itemId: number }
| { kind: "gatedGaps"; itemId: number; gaps: GatedInput[] }- The LLM selects and annotates; the developer owns the schema and the components; the engine owns the facts. Displayed numbers are the engine's, never the model's.
- Critical trick: the model returns references (ids), and the server re-hydrates the actual data (e.g.
rows) from the tool result it already computed — so the model can never drift a number. - Fits the existing pipeline: Zod →
nestjs-zodDTO → OpenAPI → Orval client. The same Zod schema (viazod-to-json-schema) becomes the model's structured-output schema — one source of truth for "what shapes can come back."
Locked early (load-bearing, painful to change)
- Endpoint shape: free-text in, typed
blocks[]out. The FE binds to this and nothing else. - Capability layer as a clean typed boundary — each capability a typed function, logic in services, not controllers. This becomes both the router target and the agent tool later.
- We own the block/entity union (starts small).
Deferred (swappable per stage, no rewrite)
- Router (stage 1–2) vs agent loop (stage 3); model/provider;
@anthropic-ai/sdkvs Vercel AI SDK (note: Vercel'sstreamUIis RSC/Next-shaped and does not fit our Vite SPA — we can use the pattern and@anthropic-ai/sdkdirectly, which the brief already commits to); streaming; narration templated vs LLM; number of intents/tools.
Parked for their own focused talks
- Structured LLM replies, in depth (schema design, validation, streaming).
- FE/BE "build pages structurally" — the generative-UI rendering contract.
- Design-system + "lego" component approach for entity-typed rendering.
- Ontology / knowledge graph. GW2 is a strong candidate (rich, stable entity relations). The official GW2 Wiki reportedly runs Semantic MediaWiki
[NEEDS VERIFICATION]— part of the knowledge half may already be queryable upstream. Spectrum: bag-of-tools → typed entity/relation layer (the sweet spot; ~80% of the value, no OWL/RDF/reasoner) → full formal ontology (a real learning target, but a whole subsystem). Ontology helps most for the knowledge questions (GraphRAG), least for the computation ones the recipe-graph engine already nails.
Existing assets this builds on
apps/api/src/legendaries/ranking.service.ts—rank(apiKey) → RankingRow[](the Stage-1 capability).apps/api/src/legendaries/ranking.schema.ts—RankingRow,GatedInput(Zod, drives OpenAPI/Orval).apps/api/src/recipe-graph/recipe-graph.service.ts—resolvePriced(itemId)(buy-vs-craft tree).apps/api/src/account/account.service.ts— materials/wallet/owned items (needsX-GW2-Key).apps/api/src/mcp/mcp.tools.ts— the tool surface (spec 017) that becomes the agent's tools later.
Open questions to resolve in the spec
- Exact scope of 029: seam + exactly one capability (Stage 1), or a small menu (Stage 2)?
- What "closest to craft" means precisely — the scoring (gated-inputs-satisfied vs remaining gold vs a blend), and how to rank rows that are already
craftable. - Narration: templated or a second LLM call, for v1.
- How the free-text endpoint receives the account key (reuse the
X-GW2-Keyheader convention). - How to test a non-deterministic front door (eval the parse; unit-test the capability + block assembly).
- LLM provider/model + cost/latency guardrails (timeout, max output).