Spec 017 — MCP server over the GW2 API and our own computed surface
Status: implemented Branch: 017-mcp-server
Status is set by the human, never by the agent. It moves draft → approved → implemented.
Amendment A (2026-08-18) — APPROVED. Everything marked [A] below was added after the original spec was approved and its four tools were implemented. The human approved the amendment on 2026-08-18; the agent transcribed that decision here on explicit instruction and did not set it on its own initiative. The
approvedstatus above now covers the original scope and Amendment A.Amendment A adds six
priory_*tools over this project's own computed endpoints, including four that read the caller's GW2 account. It arrives because seven specs (018–023) merged while 017 was in review, and the surface they built is exactly what an agent needs and cannot reach.Revision A2 (2026-08-18) — APPROVED. Discovery answered all four of Amendment A's markers and refuted two of them (A2, A3). Per the constitution a refuted claim sends the spec back to step 1, so R13, R16, SC20 were revised, R18 and SC21 added, and one new marker (A5) opened — since discharged from code. The human approved this revision on 2026-08-18; the agent transcribed that decision on explicit instruction and did not set it on its own initiative.
All five Amendment A markers now have verdicts in
research.md, which iscomplete. The plan gate is open.What changed, in one line each:
priory_account_materialscalled with noidsnow returns a per-category roll-up instead of 468 rows (~18,100 tokens, measured);priory_legendary_rankingstays but declares its >31 s cost by human decision; and R13's claim that the amendment adds no new logic was false and is withdrawn — the roll-up is new logic, confined to a pure shaping function.
Problem
Agents working on this project — Claude Code during spec, research and review — constantly need facts from the GW2 API v2: what item 19721 is, what 30703 costs on the Trading Post, which recipes output a given item. Today they get those facts by hand-curling api.guildwars2.com, which has three costs. The agent must already know the v2 surface (there is no discovery, so it guesses paths and query shapes and retries); the raw payloads are fat with icon, description, chat_link and nested details that no decision depends on, burning context on every lookup; and the agent operates outside the rate-limit budgeting, batching and caching that Gw2Service already implements, so it re-fetches what we hold and can trip 429 on its own.
Meanwhile Gw2Service — items, recipes, prices, searchRecipes, behind a 300/min token bucket, a 199-id batcher and a bounded cache — is reachable only from inside apps/api. It is exactly the surface an agent wants, and nothing can call it.
Separately, this project's stated goals include exploring AI-specific topics with the same rigour as the rest of the stack. MCP is the standard answer to "expose capability to an out-of-process agent", and this is a real, small, honest use of it rather than a contrived one.
[A] The original spec wrapped the upstream GW2 API and deliberately stopped there. Since then specs 018–023 merged, and apps/api now computes things the GW2 API cannot answer at all: a priced recipe tree with buy-versus-craft decisions (/recipe-graph/:itemId), a curated legendary list carrying a generation field that exists nowhere upstream (/legendaries), the player's own holdings enriched with category and sell price (/account/materials, /account/wallet), and a personalised profit ranking that sums material storage, bank, shared inventory and every character's bags (/legendaries/ranking).
An agent can reach none of it. The original spec argued these endpoints were covered because "they already have an OpenAPI document an agent can read" — that has proved insufficient in practice, which is the revisit condition the spec itself named. An OpenAPI document tells an agent a route exists; it does not put the result in the conversation, does not shape it for a context window, and does not carry the credential. So the agent falls back to composing gw2_* calls by hand and re-deriving arithmetic the server already does correctly — which is both slower and a second implementation that can disagree with the first.
Four of those endpoints read the player's own account, which is why they were deferred: the player's GW2 key and this server's own bearer token both want the Authorization header on the same request. That is the design question Amendment A answers rather than defers.
User stories
Ordered by priority. Each story must be independently testable and shippable — if only P1 ships, there is still something usable.
P1 — An agent looks up GW2 data through typed, discoverable tools
As the developer's agent, I want to call GW2 item, recipe, price and recipe-search lookups as named MCP tools with declared input schemas and trimmed responses, so that I discover the surface instead of guessing it, spend context on facts rather than payload padding, and inherit the batching, caching and rate-limit budget the app already has.
Independent test: with the API running locally and an MCP client pointed at POST /api/mcp with a valid bearer token, tools/list returns the four tools with their descriptions and JSON Schemas, and tools/call on gw2_items with {"ids":[19721]} returns the trimmed record for Glob of Ectoplasm. No other story implemented.
Acceptance scenarios
- Given a connected MCP client, when it sends
tools/list, then the response names exactlygw2_items,gw2_recipes,gw2_prices,gw2_recipe_search, each with a human-readable description and an input schema declaring its parameters. - Given a connected client, when it calls
gw2_itemswith a list of ids, then the result contains one record per found id carrying only the declared fields, and noicon,description,chat_linkordetails. - Given a connected client, when it calls
gw2_priceswith ids, then each result carries the item id and its best buy and sell unit prices. - Given a connected client, when it calls
gw2_recipe_searchwith aninputor anoutputitem id, then the result is the list of matching recipe ids. - Given two successive calls for the same ids within the cache TTL, when the second is served, then it is answered from the existing
Gw2Servicecache without a second upstream request. - Given an upstream failure (
429, or ids that resolve to nothing), when a tool is called, then the client receives an MCP tool error whose text states what went wrong and what to do about it, and no stack trace or internal path crosses the wire.
P2 — The server is reachable on the deploy, and only by us
As the developer, I want the MCP endpoint served by the existing deployed API behind a static bearer token, so that I can point an agent at it from anywhere, while a stranger who finds the URL cannot drain our shared GW2 rate-limit budget or our Render free tier.
Independent test: against the deployed API, a POST /api/mcp carrying the correct Authorization: Bearer header completes an MCP initialize + tools/list; the same request without the header, or with a wrong token, returns 401 and reaches no tool.
Acceptance scenarios
- Given a request to
/api/mcpwith noAuthorizationheader, when it is handled, then the response is401and no tool handler runs. - Given a request with an
Authorization: Bearervalue that does not match the configured token, when it is handled, then the response is401. - Given the configured token is absent from the environment at boot, when the API starts, then it fails loudly rather than starting with an unguarded or permanently-closed endpoint.
- Given the committed
render.yaml, when it is inspected, then the API service declares the MCP token as a dashboard-set secret (not a committed value).
P3 [A] — An agent reaches this project's computed answers, not just upstream facts
As the developer's agent, I want the priced recipe tree and the curated legendary list as MCP tools whose names and descriptions say the answer is ours rather than ArenaNet's, so that I use the server's arithmetic instead of re-deriving it from raw lookups, and so that I never present a Priory computation as an official figure.
Independent test: with a connected client, tools/list returns ten tools; the four gw2_* describe themselves as official-API pass-throughs and the six priory_* as Priory-computed; tools/call on priory_recipe_tree with a craftable item id returns a priced tree, and on priory_legendaries returns the curated list including generation, which no gw2_* tool can produce.
Acceptance scenarios
- Given a connected client, when it sends
tools/list, then every tool name begins with eithergw2_orpriory_, and every description's first clause states which of the two it is. - Given a connected client, when it calls
priory_recipe_treewith a craftable item id, then the result is the priced tree — the same computationGET /api/recipe-graph/:itemIdserves, not an unpriced one. - Given an item id the GW2 API does not know, when
priory_recipe_treeis called with it, then the client receives a tool error saying the item was not found, matching the404the REST route returns for the same id. - Given a connected client, when it calls
priory_legendaries, then each row carriesgeneration, and the tool description states thatgenerationis this project's curation and not an upstream field.
P4 [A] — An agent answers questions about my account, without my key entering the conversation
As a player whose agent is connected to this server, I want account-scoped tools that read my own holdings, wallet and profit ranking, with my GW2 key travelling as client configuration rather than as something the model types, so that I get personalised answers without my credential being written into a transcript, a tool-call log, or anything that ingests either.
Independent test: a client configured with both an Authorization bearer and an X-GW2-Key header calls priory_account_materials with no key argument and receives the caller's own owned materials; inspecting the request the model made shows the tool arguments contain no credential.
Acceptance scenarios
- Given a client configured with a valid
X-GW2-Keyheader, when it callspriory_account_materialswith no arguments, then the result is that account's owned materials. - Given the same client, when any account tool is called, then the tool's declared input schema contains no key, token or credential parameter — the credential cannot be passed as an argument even by a model that tries.
- Given a client with no
X-GW2-Keyheader, when it calls an account tool, then the tool is still listed bytools/list, and the call returns an error naming the header to configure — not a generic failure and not a missing-tool answer. - Given a client whose
X-GW2-Keyis invalid or expired, when it calls an account tool, then the error says the key was rejected, distinctly from the missing-header case. - Given a key lacking a scope the endpoint needs, when an account tool is called, then the error names the missing scope.
- Given any account-tool call, successful or failed, when the response and any log line are inspected, then the key value appears in neither.
Requirements
R1 — A
McpModuleatapps/api/src/mcp/following the feature-module layout indocs/architecture/nestjs.md: a thin controller that only hands the request to the transport, the server/tool wiring beside it, and a guard. It importsGw2Moduleand depends onGw2Service— it does not construct its ownGw2Client, so the existing token bucket, 199-id batching and bounded cache are shared with the app's own traffic (that sharing is the intended behaviour, not a compromise).R2 — Transport is MCP Streamable HTTP, served at
POST /api/mcp(under the global/apiprefix from spec 015), targeting protocol revision2025-11-25— the latest revision@modelcontextprotocol/sdk1.30.0 implements, and the one Claude Code 2.1.232 negotiates (research F1). The newer2026-07-28revision is explicitly not targeted: no published SDK implements it, and building to it by hand is rejected under Assumptions. Its obligations are recorded in R10 as a migration checklist, not as requirements here.The server runs stateless, for two independent reasons. It survives Render's single sleeping instance and any future horizontal scaling without session affinity; and the SDK's stateless transport cannot be reused across requests — it throws on the second — so a fresh MCP server and a fresh transport are constructed per request. That is a hard constraint of the library, not a style choice, and any implementation that caches either object is wrong.
GETandDELETEon the same path return405. Neither is needed: stateless means no session id, so no client ever issuesDELETE, and the oneGETa client sends (a standalone-SSE probe) is tolerated silently whether it is answered404or405.405is the documented courtesy, and it is what the newer revision will require.R3 — Access is a static bearer token read from a single environment variable. A request whose
Authorizationheader is missing or does not match returns401before any tool handler runs. Comparison is constant-time. The variable being unset at boot fails loudly (throws during bootstrap), following theVITE_APP_ENVprecedent in spec 015 R4 — an unguarded MCP endpoint and a silently-dead one are both worse than a refused start. The token value is never logged.The refusal must not look OAuth-capable. A rejected request returns a bare
401with noWWW-Authenticateheader, and the service exposes no/.well-known/oauth-protected-resourceand no/.well-known/oauth-authorization-server. A client that sees any OAuth hint may abandon the configured bearer header and pursue an authorization-server discovery flow instead of reporting a failed connection (research V4). The401body carries a short readable message, since the client surfaces it. This is a requirement, not a nicety: it is what keeps a static token workable.Originis validated as a second, independent check. The MCP specification makes this a MUST on all incoming connections, to prevent DNS-rebinding attacks in which a page in the developer's browser reaches a locally-bound MCP server. AnOriginheader that is present and not in the allow-list is refused with403, before authentication and before any tool handler runs. The allow-list is the existingCORS_ALLOWLIST(apps/api/src/config/cors.ts, spec 015 R2), so there is one list, not two that can drift.A request with no
Originheader is allowed through to the bearer check. This is the normal case and must not be broken: a native MCP client is not a browser and sends noOrigin— the wire log in research F1 confirms Claude Code sends none. Refusing an absentOriginwould reject every real client while stopping no attack, since the browser is precisely the agent that always sets it.R4 — Exactly four tools, each a thin wrapper over an existing
Gw2Servicemethod, with a description written for a model rather than a human reader, and a Zod input schema surfaced as JSON Schema:Tool Input Result fields gw2_itemsids: number[]id,name,rarity,type,flags,vendor_valuegw2_recipesids: number[]id,output_item_id,output_item_count,disciplines,min_rating,ingredientsgw2_pricesids: number[]id,buy(best buy unit price),sell(best sell unit price)gw2_recipe_searchinputoroutputitem id (exactly one)recipe ids R5 — Tool results carry only the fields in R4; every other upstream field is dropped at the tool boundary. The shaping is asserted by test, so a future upstream field cannot silently re-enter the payload. Measured on 50 representative ids (research V3), against minified upstream JSON — not the pretty-printed bytes the API actually sends, whose whitespace alone is 39% of the items payload and would flatter the number:
Tool reduction gw2_items72.9% (28,137 B → 7,623 B) gw2_prices72.7% (6,857 B → 1,872 B) gw2_recipes31.3% (16,617 B → 11,415 B) The weight is in
details(32.3%),icon(17.0%) andgame_types(8.7%) — notdescription(5.8%) orchat_link(4.6%), which an earlier draft of this requirement wrongly blamed.gw2_recipesis the weak case and is stated as such: its shape keepsingredientsanddisciplines, which are most of a recipe record.This is a context-window saving, not a bandwidth one — HTTP gzips the wire regardless. The criterion that matters is tokens spent per lookup, not bytes transferred.
flagsis retained deliberately despite being 31% of the shaped items payload — the most expensive field kept — because buyable-versus-gated classification depends on it.levelis not among the item fields.Gw2ItemSchemadoes not declare it andz.objectstrips unknown keys, so exposing it would mean widening the GW2 client's schema for the lowest-value field in the set. Dropped rather than paid for.R6 — Failures map to MCP tool errors with actionable text, reusing the existing
gw2.errors.tstaxonomy: rate-limit exhaustion says so and says to retry, a lookup that resolves to nothing says which ids were not found, and an unexpected upstream failure is reported without leaking stack traces, internal paths or the upstream URL.R7 — Ids are validated at the tool boundary before reaching
Gw2Service: positive integers, and a bounded list length so one call cannot fan out into an unbounded number of upstream batches. The bound is a named constant, stated in the tool description so the model can respect it.R8 — The MCP route is excluded from
openapi.json(@ApiExcludeEndpoint). It is JSON-RPC over a single POST, not a REST resource; including it would make orval generate a meaningless client hook inapps/weband would churn the committed contract.verify:contracttherefore stays green withopenapi.jsonunchanged. The endpoint is specified here and in R10 instead — the exclusion is a deliberate decision, not a documentation gap.R9 —
render.yamldeclares the MCP token on the API service as a dashboard-set secret (sync: false), never a committed value, andtests/deploy/render-blueprint.test.tsasserts it.R10 —
docs/architecture/mcp.mdrecords the tool contract, the transport and stateless decision with its reasoning, the auth posture and its limits, a copy-pasteable client configuration for both the local and the deployed server, and the2026-07-28migration checklist (mandatoryMcp-Method/Mcp-Nameheaders with400+-32020validation,server/discover,ttlMs/cacheScopeontools/list,202on notifications,404+-32601on unknown method, and the reconnect hazard of advertisingsubscriptions/listenon a host that cuts long-held streams).The configuration uses the documented
type/url/headersshape — noting that aurlentry withouttypeis a configuration error, read as a stdio server. The token must not be committed.${VAR}expansion insideheadersis documented but has a live history of not substituting, whose failure mode is that the literal${VAR}string is sent and produces a confusing401(research V4). Implementation therefore verifies expansion empirically and, if it does not work, falls back to the documentedheadersHelperor to local-scope configuration in~/.claude.json. Whichever lands, the committed repository contains no token.R11 — No artifact under any
docs/superpowers/path; prior specs' suites still pass, updated only where 017 changes their subject.R12 [A] — Tool names carry provenance. A
gw2_prefix means a pass-through to the official GW2 API: shaped, but not interpreted. Apriory_prefix means this project computed the answer. Every description's first clause states which it is —Official GW2 API: …orPriory-computed: …— because that string is the only thing a model reads when choosing a tool. Apriory_description also names what makes the answer ours, so the model can attribute it correctly rather than reporting our arithmetic as an ArenaNet figure. This is asserted by test over every registered tool, not left to authorial discipline.R13 [A] — Six further tools, each a wrapper over a service that already exists and is already tested:
Tool Key? Input Result priory_recipe_treeno itemId(positive int)the priced buy-versus-craft tree priory_legendariesno optional generation, optionaltypecurated legendaries incl. generationpriory_accountyes none the account display name priory_account_materialsyes optional bounded idsper-category roll-up, or those ids (R16) priory_account_walletyes none wallet entries priory_legendary_rankingyes none personalised profit ranking — slow, see R18 One piece of new logic, named rather than smuggled. An earlier draft of this requirement claimed the amendment "introduces no new domain logic". R16's per-category roll-up is new logic — a grouping and three sums — so the claim was false and is withdrawn. The roll-up is confined to a pure function in the MCP shaping layer over
AccountService.getMaterials's existing output. No service changes, nothing for the web UI to inherit, and it is unit-testable without a key. Everything else in this table remains a thin wrapper.Each calls its Nest service directly. No tool issues an HTTP request to this API's own REST routes: that would be a round trip out through our own guard and back, and would make the tool's behaviour depend on the deploy's own reachability.
R14 [A] — The player's GW2 key travels as a separate request header,
X-GW2-Key, set once in the client's MCP configuration beside theAuthorizationheader. Consequences, each load-bearing:- The guard is unchanged. It still authenticates only
MCP_AUTH_TOKEN. The player key is not an authentication input to us — it is a credential we forward to GW2, exactly as the REST controllers inaccount.controller.tsandlegendaries.controller.tsalready do. A malformed key is therefore not a403from the guard; it is a rejected upstream call surfaced as a tool error. - No account tool declares a key parameter. The credential cannot arrive as a tool argument even if a model attempts it, which is what keeps it out of transcripts and tool-call logs.
- The key is read per request and passed into the per-request server construction. The stateless transport already forces a fresh server per request (R2), so this needs no new lifecycle.
- The key is never logged, never persisted and never returned.
apps/apistores no key today (AccountServiceis pass-through) and this amendment does not change that. - A single shared
MCP_AUTH_TOKENstill gates access;X-GW2-Keydecides whose account is read. The two are independent, and the limits of the shared token in R3 are unchanged by this. - R10's documented client configuration gains the second header, with the same warning already recorded there:
${VAR}expansion insideheadershas a live history of not substituting, and its failure mode here is worse than a confusing401— an unexpandedX-GW2-Keyreaches GW2 as a literal string and is rejected as an invalid key, which reads as "my key is wrong" rather than "my config did not expand".docs/architecture/mcp.mdstates this explicitly.
- The guard is unchanged. It still authenticates only
R15 [A] — Three further error cases, mapped to distinct, actionable text.
toToolErrorTextcurrently falls throughGw2UnauthorizedErrorandGw2ForbiddenErrorto "failed for an unexpected reason", which is useless to a caller who can fix the problem:Case Text names No X-GW2-Keyheaderthe header to add to the MCP client configuration Key invalid or expired that the key was rejected, and to re-issue it Key missing a scope the missing scope (public information, carried by Gw2ForbiddenError.scope)The three stay distinguishable: collapsing them into one message reproduces exactly the dead end this requirement exists to remove. No message includes the key value.
Scope naming applies to three tools, not four (research A4).
priory_account_materials,priory_account_walletandpriory_legendary_rankingreachauthedCachedRead, which maps403toGw2ForbiddenErrorcarrying a parsed or fallback scope.priory_accountdoes not:Gw2Client.account()maps401and403alike toGw2UnauthorizedError, so it can only report a rejected key. That costs nothing real —/v2/accountneeds theaccountscope, which is mandatory on every GW2 key and cannot be missing — andGw2Clientis not changed to chase it. The exception is stated here rather than left as a rule one tool silently fails to honour.R16 [A] — The new tools are shaped like the old ones (R5), by the same rule and for the same reason — context window, not bandwidth:
iconis dropped from every result. An agent cannot render a URL, and it is pure padding.priory_account_materialsaccepts an optional boundedidsarray — the same bound and the same boundary validation as R7 — which returns exactly those ids includingcount: 0, since "I own none of this" is the answer to "do I have enough". This is the shape that pairs withpriory_recipe_tree, and it is unchanged by the revision below.- Called with no
ids, it returns a per-category roll-up — not rows. One entry per material category carrying the category name, the number of distinct owned items, the total item count and the total sell value. Roughly nine entries. - Absence of
iconis asserted by test, so a future upstream field cannot re-enter the payload.
Revision, 2026-08-18 — the previous default was refuted, not adjusted. This requirement first said the no-
idscall returns every owned row (count > 0), on the assumption that unowned slots dominate. Measured against a real account (research A2), they do not: 69% of rows are owned, so the filter removed less than a third of them and the shaped result was 468 rows, 70.6 KB, ~18,100 tokens — a 38.3% reduction where an order of magnitude was assumed. That is roughly 2.4× the entire measured saving of the original four tools (R5) spent on one call, and it would evict most of what the calling agent was holding. A tool that costs more context than it saves defeats the purpose of the whole feature, so the default changed shape rather than being tuned.The roll-up answers "what is in my storage" at the altitude that question is actually asked, in roughly 2% of the tokens, and the
idsform answers "do I have enough of these" precisely. Neither question needs 468 rows. No row-level browse is offered at all — an agent that wants rows names ids.R18 [A] —
priory_legendary_rankingships with its cost declared, and the declaration is not a fix. Research A3 measured 22 upstream requests and 31.4 seconds for thegetOwnedItemsfan-out alone on a 19-character account; the fullrank()adds recipe-graph resolution and pricing on top, so ">31 s" is the floor. Human decision, 2026-08-18: keep the tool, declare the cost.What that means concretely, and what it does not:
- The description states the cost — that the call takes tens of seconds, scales with character count, and consumes a meaningful share of a shared rate-limit budget — so a model can decide not to call it. This is the whole benefit.
- It does not prevent timeouts. A call exceeding the client's tool timeout fails. The declaration means the failure is expected rather than mysterious; it does not make the call succeed. The spec says so plainly instead of implying the caveat solves the problem.
- The description advises one retry.
ACCOUNT_TTL_MSis 5 minutes (gw2-client.ts:52) and noAbortSignalis passed to the upstream fetches, so a client disconnect should not cancel work already in flight — a first call that times out is expected to leave the account reads cached, making a retry within five minutes fast. This is inferred from the code, not measured (A5 (discharged)), and the description must not promise it as certain. - The ceiling is recorded in code, as a
ponytail:comment naming what it is (cost linear in character count, not tunable) so it lands in the debt ledger rather than being rediscovered. - No optimisation is designed here. There is no agreed approach, and inventing one to justify keeping the tool would be worse than shipping it honestly slow.
R17 [A] — Two changes outside
apps/api/src/mcp/, both required and both minimal:LegendariesModuledeclares noexports, so neitherLegendariesServicenorRankingServicecan be injected anywhere else. Both are exported soMcpModulecan depend on them. Nothing else changes in that module.- The pricing step of
GET /recipe-graph/:itemIdlives in the controller (collectPricedIds→gw2.prices→priceGraph), so a tool callingRecipeGraphService.resolve()alone would get an unpriced tree. That step moves into the service, and the controller calls it — one priced-tree implementation with two callers, rather than two that can silently diverge. The route's observable behaviour, including its404, is unchanged, and its existing tests must pass untouched to prove it.
buildToolstakes a dependencies object rather than positional parameters at this point: five services plus an optional per-request credential is past the threshold where positional arguments stay readable.
Mark anything unresolved inline rather than assuming an answer. Two markers, split by who can answer:
[NEEDS CLARIFICATION: specific question]— only the human can answer. Blocks step 1.5.[NEEDS VERIFICATION: specific question]— only reality can answer, inresearch.mdwith cited evidence. Blocks the approval gate.
No [NEEDS CLARIFICATION] remains — the surface (the four endpoints we already wrap), the transport (streamable HTTP rather than stdio, chosen for learning and professional fit), the auth posture (static bearer over OAuth) and the protocol depth (tools only, no resources or prompts) were all decided in brainstorming.
No [NEEDS VERIFICATION] remains either. All five markers have verdicts in research.md, which is complete:
| Was | Verdict |
|---|---|
| Is Streamable HTTP current, HTTP+SSE deprecated, stateless legal? | Confirmed (V1) — and the spec has since moved to 2026-07-28, which no SDK implements; R2 targets 2025-11-25 accordingly. |
| Does the SDK transport work behind Nest's Fastify adapter, stateless? | Confirmed (V2) — proven by spike; the glue is passing req.body to handleRequest. |
| What is the measured payload reduction? | Confirmed (V3) — 72.9% / 72.7% / 31.3%, folded into R5. |
| How does a client attach a static bearer header without committing it? | Confirmed with constraints (V4) — folded into R3 and R10. |
| Is a scoped name→id index cheap enough? | Premise refuted (V5) — 2 requests, not 350. Excluded from 017 on scope grounds by human decision, not on cost. |
One finding (F1) is load-bearing enough to name here: the published SDK does not implement the current spec revision, and neither does the tested client — they interoperate on 2025-11-25, verified with a live tools/call. R2 and R10 carry the consequence.
[A] Amendment A reopened discovery, and two of its claims were refuted. The statement above describes the original scope and remains true of it. Amendment A's four markers now all have verdicts in research.md:
| Marker | Verdict | What it did to this spec |
|---|---|---|
| A1 — do configured headers reach every request? | Confirmed — 6/6 and 11/11 on a live wire log | R14 stands unchanged |
A4 — is Gw2ForbiddenError.scope usable? | Confirmed for 3 of 4 tools | R15 and SC18 narrowed |
| A2 — is a shaped materials result small enough? | REFUTED — 468 rows, ~18,100 tokens | R16's default replaced |
| A3 — is the ranking tool affordable? | REFUTED — >31 s, 22 requests | R18 added; tool kept with declared cost |
A5 (discharged) — discharged. Does a client-side timeout leave the account reads cached, so R18's advised retry is fast? Confirmed from code (research A5): no AbortSignal exists anywhere in apps/api, the inbound and outbound sockets are unrelated, a JS handler promise cannot be cancelled, and authedCachedRead writes the cache before returning. One constraint falls out of it: accountCache is in-process, so a Render spin-down empties it — R18's description says a retry is likely fast, never that it is fast. No [NEEDS VERIFICATION] remains open on Amendment A.
The original four markers' text is kept below for the record:
- A1 (discharged) — Does an MCP client actually forward a second custom header (
X-GW2-Key) on everytools/call, or only on initialize? R14 collapses if the header does not reach each call. Research V4 already found${VAR}expansion inheadersunreliable (R10); this needs the same empirical treatment, on a real client, with a wire log. - A2 (discharged) — What is the real size of a shaped
priory_account_materialsresult on an established account: how many rows survive thecount > 0filter, and how many tokens is that? R16 assumes dropping zeroes andiconis sufficient. If an owned-materials list is still thousands of rows, the default needs a bound rather than a filter, and R16 is wrong as written. - A3 (discharged) — How long does
priory_legendary_rankingtake, and how many upstream requests does it issue?RankingService.rankreachesgetOwnedItems, which fans out over material storage, bank, shared inventory and one request per character. Behind a single tool call, on a shared 300/min budget and a free-tier instance, this may be the heaviest thing in the codebase. If it is too slow or too costly, it does not ship as a tool in this amendment. - A4 (discharged) — Does
Gw2ForbiddenError.scopeactually carry a usable scope name for the account endpoints in scope, so R15's "name the missing scope" is implementable? The 016 → 018 correction recorded indocs/gaps/gw2-auth-status.mdshows this project has already been wrong once about GW2's auth-failure semantics; it is not assumed a second time.
No [NEEDS CLARIFICATION] is introduced: the key transport, the absent-header behaviour and the materials bounding were all decided with the human in brainstorming before this amendment was written.
Success criteria
Measurable and outcome-focused. MCP, Nest and Render are named because the change is about that wiring, per Assumptions.
- SC1 — An MCP client that completes initialization against a running instance receives exactly four tools from
tools/list, each with a non-empty description and an input schema. (integration test over the SDK's in-memory transport) - SC2 — Each of the four tools, called with valid input, returns the fields declared in R4 and none of the excluded ones. (per-tool test against a mocked
Gw2Service) - SC3 — A request to
/api/mcpwithout a valid bearer token receives401and executes no tool handler. (guard test + bootstrap test against a running instance) - SC4 — Booting the API with the MCP token variable unset throws rather than starting. (automated test over the config resolver)
- SC5 — Repeated identical lookups within the cache window produce one upstream request, and no tool path constructs a
Gw2Clientof its own. (test asserting call counts through a mocked client; the architecture test inconventions/guards the second half) - SC6 — Upstream
429and not-found conditions surface as MCP tool errors whose text names the cause and the remedy, with no stack trace, internal path or upstream URL in the payload. (error-mapping tests) - SC7 — A single tool call cannot exceed the declared id-list bound; over-long or non-positive input is rejected at the boundary before any upstream call. (validation tests)
- SC8 —
openapi.jsonis unchanged andverify:contractpasses. (CI step) - SC9 — The committed
render.yamlcarries the MCP token as a dashboard-set secret with no value in the repository. (structural blueprint test) - SC10 — On the deployed API, a real MCP client authenticates, lists the tools and successfully calls one of them. (observational — verified on the first deploy, recorded with a date in
docs/architecture/mcp.md) - SC11 — The count of files under any
docs/superpowers/path stays zero, and prior specs' suites still pass. (existing invariants) - SC12 — Every acceptance scenario and success criterion maps to a named test or a dated manual record, with no gap; the automated portion passes.
- SC13 — A request carrying an
Originheader outside the allow-list is refused403before authentication and before any tool handler runs, while a request with noOriginheader proceeds to the bearer check and, with a valid token, succeeds. (guard tests, both directions) - SC14 [A] —
tools/listreturns exactly ten tools; every name begins withgw2_orpriory_, and every description's first clause declares which. (test iterating every registered tool, so a future tool cannot be added without a provenance clause) - SC15 [A] —
priory_recipe_treereturns the priced tree, and the value it returns for a given id is the same oneGET /api/recipe-graph/:itemIdreturns for that id. (test over the shared service; the existing recipe-graph route tests pass untouched, proving the extraction in R17 changed no behaviour) - SC16 [A] — No account tool's input schema contains a key, token or credential parameter. (schema-level test over every registered tool, asserting absence by property name)
- SC17 [A] — An account tool called with no
X-GW2-Keyheader is still listed, and returns an error naming the header to configure. (tool test; distinct from the invalid-key case) - SC18 [A] — An invalid key and a key missing a scope produce different errors, the latter naming the scope, and neither includes the key value. Asserted for the three tools that can distinguish the cases;
priory_accountasserts only the invalid-key message, per R15 and research A4. (error-mapping tests, one per case) - SC19 [A] — The key value appears in no tool result and in no log line emitted during an account tool call. (test asserting absence across the response payload and captured log output)
- SC20 [A] —
priory_account_materialswith no arguments returns a per-category roll-up of roughly nine entries, containing no item rows and noicon; called with anidslist it returns exactly those ids withcount: 0included, and rejects a list over the bound before any upstream call. (shaping and validation tests over the pure roll-up function, no key required) - SC21 [A] —
priory_legendary_ranking's description states its latency, that the cost scales with character count, and the retry advice from R18. (description assertion, alongside SC14's provenance-clause check)
Out of scope
Item search by name — deferred to its own spec, and no longer on cost grounds. The original reason (roughly 350 batched requests to index all ~70k items) was refuted in discovery: an index scoped to what this project already covers is 263 distinct ids — 2 batched requests (research V5). It is excluded because it is not a thin wrapper like the other four. It needs a boot hook, and
apps/apihas none; it needs an answer for Render's free-tier spin-down emptying every cache, since an index built from live cache contents is empty exactly when an agent first asks; and it covers gen-1 trees only, 145 of 198 legendaries being single-node until a curation spec deepens them. That is a feature with its own design questions, not a fifth tool. 017 ships four tools.Authenticated GW2 account endpoints— superseded by Amendment A. The exclusion below is left as written rather than deleted: it correctly identified theAuthorization-collision problem and correctly refused to guess an answer. Amendment A answers it (R14, a secondX-GW2-Keyheader) instead of deferring it, so/account,/account/materials,/account/walletand/legendaries/rankingare now in scope./bankstays out — it has no REST route of its own; it is an input togetOwnedItems, not an endpoint. Original text:Authenticated GW2 account endpoints (
/account/materials,/wallet,/bank). Note what is already true: spec 016 is merged (3a77065), soapps/api/src/account/exists,GET /api/accountaccepts a player key viaAuthorization: Bearer <key>and validates it, andGw2Service.account()is available. Intake is therefore not the blocker. These stay out of 017 for a different reason: an account tool is not a thin wrapper. The player's GW2 key and this server's own bearer token would both want theAuthorizationheader on the same request, so the key has to travel some other way — a second header, a tool argument, or server-side configuration — and each choice has different exposure. That is design work with its own spec. 016 stores nothing (AccountServiceis pass-through), so no key-storage question is settled either.The front-end prompt → backend agent feature. A user typing a goal and getting a tailored reply needs an LLM host loop, a chat endpoint, streaming and cost control. It is the eventual destination and it is a separate spec; note that the loop, living in-process, will call the tool layer directly rather than through MCP.
Our own computed endpoints as tools— superseded by Amendment A, on the revisit condition this entry itself named. Original text:Our own computed endpoints as tools (
/legendaries,/recipe-graph/:itemId). They already have an OpenAPI document an agent can read. Revisit only if that proves insufficient in practice.It proved insufficient, for a reason the entry did not anticipate: an OpenAPI document tells an agent a route exists, but does not put its result in the conversation, does not shape it for a context window, and does not carry the credential. The agent's fallback is to re-derive the computation from
gw2_*calls — a second implementation that can disagree with the server's./healthand/commerce/pricesas tools [A]./healthanswers nothing an agent would ask./commerce/pricesisGw2Service.prices()with one field dropped (commerce.service.ts), whichgw2_pricesalready answers in a flatter shape — a second tool for the same question makes tool selection worse, not better. Ten tools, not twelve.MCP resources and prompts. Tools are the part agents use and the part that carries the design lessons.
OAuth 2.1 per the MCP authorization spec. The right answer for a public server and a substantial feature in its own right; a static token is where internal servers legitimately start.
Stateful sessions, server→client notifications, sampling, progress reporting. Nothing here is long-running enough to need them.
Rate limiting or quotas on the MCP endpoint itself. The bearer token bounds who can call it; the existing token bucket bounds what reaches upstream.
Assumptions
Gw2Serviceis the right seam.items,recipes,pricesandsearchRecipesexist, are tested, and already sit behind the token bucket, the 199-id batcher and the bounded cache. This spec adds a protocol surface over them and changes none of their behaviour.- The global
/apiprefix and the Fastify adapter stand (spec 015): the MCP route lives under the prefix, and any transport integration must work through@nestjs/platform-fastify, not Express. - The GW2 endpoints in scope need no player authentication. Items, recipes, prices and recipe search are public; the only secret this feature introduces is our own bearer token.
- The deploy pipeline stands (specs 014/015): the API is a Render service deploying on
checksPass, so the MCP endpoint ships with it and needs one new environment variable, not new infrastructure. @modelcontextprotocol/sdk(1.30.0) is a new dependency and is accepted as one — implementing the protocol by hand would be re-writing a specification, not saving weight. Its cost is known and accepted: it carries hard, non-optional dependencies onexpress,hono,@hono/node-server,cors,jose,ajv,eventsourceandexpress-rate-limit— real install weight for two classes, and it pulls in Express, which this project deliberately does not use. Nothing inapps/apiimports any of them; they arrive transitively.- A static bearer token is proportionate. The data exposed is public GW2 information; the asset being protected is our rate-limit budget and free-tier compute, not the data.
- [A] The five services are the right seam, for the same reason
Gw2Servicewas.RecipeGraphService,LegendariesService,RankingServiceandAccountServiceexist, are tested, and already sit above the shared token bucket and cache. Amendment A adds a protocol surface over them and changes their behaviour nowhere — with the single exception of R17's pricing extraction, which moves code without changing what the route returns. - [A] The account tools are only as private as the client's config file.
X-GW2-Keylives in~/.claude.json(local scope) in plaintext, which is the same posture spec 016 already accepted forlocalStoragein the browser. This amendment does not improve on it and does not pretend to: it moves the key out of the conversation, which is a different and narrower claim. A key in a local config file readable by the user's own processes is a known, accepted limitation; a key in a transcript is not, because transcripts travel. - [A] The
MCP_AUTH_TOKENholder is trusted to hold a GW2 key too. Anyone who can call the account tools is someone you gave the MCP token to, and they supply their own GW2 key — so a shared MCP token does not expose anyone's account to anyone else. This is the one place the shared-token limitation in R3 does not bite, and it is why account tools are viable before OAuth exists.
Traceability
Each acceptance scenario and success criterion maps to a named test or a dated manual-verification record — the latter only for the criterion that is inherently observational (a real client against the real deploy). SC12 asserts no empty cell. Filled in during implementation.
[A] Amendment A's rows (P3 #1–4, P4 #1–6, SC14–SC20) are listed below as PENDING and are filled in when the amendment is implemented.
A hole worth naming: SC12's existing test asserts that no cell is empty and that every backtick-cited .ts path exists on disk. A cell reading PENDING satisfies both — it is non-empty and cites no path — so the placeholder would pass the check silently. Implementing this amendment therefore includes extending that assertion to reject the literal PENDING, otherwise the traceability gate cannot tell a filled row from an unfilled one, which is the one thing it exists to do.
| Criterion | Test / verification |
|---|---|
| P1 #1 | apps/api/src/mcp/mcp.server.test.ts — SC1: tools/list returns exactly four described tools; apps/api/src/mcp/mcp.controller.test.ts — P1 #1: initialize then tools/list returns the four tools |
| P1 #2 | apps/api/src/mcp/mcp.tools.test.ts — SC2: gw2_items shapes the service result; apps/api/src/mcp/mcp.shape.test.ts — SC2: shapeItem keeps the declared fields and drops the rest (asserts absence of icon and details) |
| P1 #3 | apps/api/src/mcp/mcp.shape.test.ts — SC2: shapePrice flattens buys/sells and drops whitelisted; live call recorded 2026-08-16 in docs/architecture/mcp.md (gw2_prices {"ids":[19721]} → [{"id":19721,"buy":2603,"sell":2797}]) |
| P1 #4 | apps/api/src/mcp/mcp.tools.test.ts — P1 #4: gw2_recipe_search returns matching recipe ids, by input or by output (happy path, both branches) and R7: gw2_recipe_search requires exactly one of input/output (rejection). Corroborated by a live call on 2026-08-16 recorded in docs/architecture/mcp.md: {"output":19685} → [21] |
| P1 #5 | apps/api/src/gw2/gw2-client.test.ts — SC5/P1#5: prices() for an id twice within 60 s issues one fetch; a third call after advancing past 60 s issues a second fetch (the cache the MCP path inherits); apps/api/src/mcp/mcp.tools.test.ts — SC5: one call reaches the service exactly once (no fan-out above it) |
| P1 #6 | apps/api/src/mcp/mcp.errors.test.ts — SC6: a rate-limit failure says so and says what to do, SC6: an unexpected failure leaks no stack, path or upstream URL; apps/api/src/mcp/mcp.tools.test.ts — SC6: an upstream failure surfaces as mapped text, not the raw error |
| P2 #1 | apps/api/src/mcp/mcp.controller.test.ts — SC3/R3: rejects a missing token with a BARE 401 — the response itself carries no WWW-Authenticate; apps/api/src/mcp/mcp.guard.test.ts — SC3: rejects a missing Authorization header (throws before any handler) |
| P2 #2 | apps/api/src/mcp/mcp.controller.test.ts — SC3: rejects a wrong token; apps/api/src/mcp/mcp.guard.test.ts — SC3: rejects a wrong token |
| P2 #3 | apps/api/src/mcp/mcp.config.test.ts — SC4: throws when unset, SC4: throws when blank; reached through the DI graph in apps/api/src/mcp/mcp.module.test.ts — 017 T3: builds its DI graph (the guard resolves the token in its constructor, so the graph cannot build without it) |
| P2 #4 | tests/deploy/render-blueprint.test.ts — 017 SC9: the api declares MCP_AUTH_TOKEN as a dashboard-set secret, with no value committed |
| SC1 | apps/api/src/mcp/mcp.server.test.ts — SC1: tools/list returns exactly four described tools, SC1: tools/call reaches the handler (both over InMemoryTransport) |
| SC2 | apps/api/src/mcp/mcp.shape.test.ts — all three SC2: cases; apps/api/src/mcp/mcp.tools.test.ts — SC2: gw2_items shapes the service result |
| SC3 | apps/api/src/mcp/mcp.guard.test.ts — SC3: rejects a missing Authorization header, SC3: rejects a wrong token, SC3: allows a valid token with no Origin header; apps/api/src/mcp/mcp.controller.test.ts — the two SC3 cases at HTTP level |
| SC4 | apps/api/src/mcp/mcp.config.test.ts — SC4: returns the configured token, SC4: throws when unset, SC4: throws when blank |
| SC5 | apps/api/src/mcp/mcp.tools.test.ts — SC5: one call reaches the service exactly once; apps/api/src/gw2/gw2-client.test.ts — SC5/P1#5: prices() for an id twice within 60 s issues one fetch…; second half (no Gw2Client under src/mcp/) by apps/api/src/conventions/conventions.arch.test.ts — SC4/G4: GW2 access only under gw2/, which runs G4 over the real tree |
| SC6 | apps/api/src/mcp/mcp.errors.test.ts — all four SC6: cases; apps/api/src/mcp/mcp.tools.test.ts — SC6: an upstream failure surfaces as mapped text, not the raw error |
| SC7 | apps/api/src/mcp/mcp.tools.test.ts — SC7: rejects an empty id list before calling the service, SC7: rejects more than MAX_IDS ids before calling the service, SC7: rejects non-positive ids (each asserts the service was not called) |
| SC8 | apps/api/src/generate-openapi.test.ts — 017 SC8: no MCP path is documented; pnpm verify:contract (CI step, green with openapi.json unchanged) |
| SC9 | tests/deploy/render-blueprint.test.ts — 017 SC9: the api declares MCP_AUTH_TOKEN as a dashboard-set secret, with no value committed |
| SC10 | Dated manual record, 2026-08-16 — verification log in docs/architecture/mcp.md: Claude Code 2.1.233 registered against a locally-run instance, claude mcp list → ✔ Connected, tools/call on gw2_items {"ids":[19721]} returned the shaped Glob of Ectoplasm from the live GW2 API. Against the deployed API it remains pending — 017 had not merged or deployed when this was written; re-run and add an entry after the first deploy |
| SC11 | tests/workflow/repo-invariants.test.ts — SC3: no workflow artifact is written outside specs/NNN-<slug>/ (asserts zero files under docs/superpowers/); prior suites via the full pnpm test run |
| SC12 | tests/deploy/render-blueprint.test.ts — 017 SC12: spec 017 traceability is complete and cited tests exist, asserting this table has rows, that no criterion cell is empty, and that every backtick-cited .ts path exists on disk. The 014 and 015 blocks beside it are hardcoded to their own spec paths, so 017 needed its own |
| SC13 | apps/api/src/mcp/mcp.guard.test.ts — SC13: rejects a disallowed Origin with 403, before the token is even considered, SC13: allows an allow-listed Origin; apps/api/src/mcp/mcp.controller.test.ts — SC13: a present, non-allow-listed Origin is refused 403 even with a valid token, SC13/F1: an ABSENT Origin is allowed — native MCP clients send none |
| P3 #1 | apps/api/src/mcp/mcp.tools.test.ts — 017 SC14/R12: every description opens by declaring its provenance and 017 SC14: exactly ten tools, every name prefixed gw2_ or priory_ (loops over every registered tool, so an eleventh cannot dodge them) |
| P3 #2 | apps/api/src/mcp/mcp.tools.test.ts — 017 P3 #2/SC15: priory_recipe_tree returns the priced root from the shared service (asserts identity, not deep equality); apps/api/src/recipe-graph/recipe-graph.service.test.ts — 017 A2/SC15: returns the priced root, pricing the buyable ids in ONE batched call |
| P3 #3 | apps/api/src/mcp/mcp.tools.test.ts — 017 P3 #3: an unknown item id becomes a not-found tool error, not a generic one; apps/api/src/recipe-graph/recipe-graph.service.test.ts — 017 A2/P3 #3: throws ItemNotFoundError when the GW2 API did not know the root id |
| P3 #4 | apps/api/src/mcp/mcp.tools.test.ts — 017 P3 #4: priory_legendaries passes the filter through and returns generation, 017 P3 #4: priory_legendaries rejects a generation outside 1-3 before the service |
| P4 #1 | apps/api/src/mcp/mcp.controller.test.ts — 017 P4 #1: the X-GW2-Key header value reaches the tool handler (real route, real guard); apps/api/src/mcp/mcp.tools.test.ts — 017 P4 #1: priory_account_materials with no ids returns the category roll-up |
| P4 #2 | apps/api/src/mcp/mcp.tools.test.ts — 017 SC16: no tool declares a credential parameter (reads property names via z.toJSONSchema, so the .refine()-wrapped schema is covered too) |
| P4 #3 | apps/api/src/mcp/mcp.tools.test.ts — 017 SC17: every account tool without a key returns the missing-header error, 017 SC17: a keyless caller still sees every tool listed; apps/api/src/mcp/mcp.controller.test.ts — 017 SC17: without the header the call is NOT a 401 — it is a tool error naming the header |
| P4 #4 | apps/api/src/mcp/mcp.errors.test.ts — 017 SC18: an invalid key says the key was rejected, and names no scope |
| P4 #5 | apps/api/src/mcp/mcp.errors.test.ts — 017 SC18: a missing scope names the scope, distinctly from an invalid key; apps/api/src/mcp/mcp.tools.test.ts — 017 SC18: a missing scope surfaces with the scope named |
| P4 #6 | apps/api/src/mcp/mcp.tools.test.ts — 017 SC19: the key never appears in a successful result; apps/api/src/mcp/mcp.controller.test.ts — 017 SC19: the key does not appear in the response body; apps/api/src/mcp/mcp.errors.test.ts — 017 SC19: no key-related branch emits anything key-shaped |
| SC14 | apps/api/src/mcp/mcp.tools.test.ts — 017 SC14: exactly ten tools, every name prefixed gw2_ or priory_, 017 SC14/R12: every description opens by declaring its provenance, 017 R12: a priory_ description never claims to be the official API; apps/api/src/mcp/mcp.controller.test.ts — 017 SC14: tools/list over the real route advertises all ten tools |
| SC15 | apps/api/src/mcp/mcp.tools.test.ts — 017 P3 #2/SC15: priory_recipe_tree returns the priced root from the shared service; apps/api/src/recipe-graph/recipe-graph.controller.test.ts — the three 008 T6 cases, assertions unchanged across the A2 extraction, proving the route's value did not move |
| SC16 | apps/api/src/mcp/mcp.tools.test.ts — 017 SC16: no tool declares a credential parameter |
| SC17 | apps/api/src/mcp/mcp.tools.test.ts — 017 SC17: every account tool without a key returns the missing-header error, 017 SC17: a blank key is treated as missing, not passed upstream, 017 SC17: a keyless caller still sees every tool listed; apps/api/src/mcp/mcp.errors.test.ts — 017 SC17: the missing-header text names the header to configure |
| SC18 | apps/api/src/mcp/mcp.errors.test.ts — 017 SC18: an invalid key says the key was rejected, and names no scope, 017 SC18: a missing scope names the scope, distinctly from an invalid key; apps/api/src/mcp/mcp.tools.test.ts — 017 SC18: a missing scope surfaces with the scope named |
| SC19 | apps/api/src/mcp/mcp.tools.test.ts — 017 SC19: the key never appears in a successful result; apps/api/src/mcp/mcp.controller.test.ts — 017 SC19: the key does not appear in the response body |
| SC20 | apps/api/src/mcp/mcp.shape.test.ts — 017 SC20: shapeMaterial drops icon, the numeric category and categoryOrder, 017 SC20: rollUpMaterials aggregates per category, owned rows only, 017 SC20: rollUpMaterials returns [] when nothing is owned, 017 SC20: shapeWalletEntry drops icon and order, 017 SC20: shapeRankingRow drops icon and keeps every money column; apps/api/src/mcp/mcp.tools.test.ts — 017 SC20: with ids it returns exactly those rows, count 0 included |
| SC21 | apps/api/src/mcp/mcp.tools.test.ts — 017 SC21/R18: the ranking description states its cost and the retry advice |