Skip to content

Research 017 — MCP server over the GW2 official API ​

Status: complete Step 1.5 output, written between the spec draft and the approval gate. open until every [NEEDS VERIFICATION] marker in spec.md has a verdict here; complete once they do. The original five markers have verdicts; the one finding that could have blocked the gate (F1) is closed by direct wire evidence.

Amendment A (2026-08-18) reopened this document. It adds markers A1–A4. All four now have verdicts — but two are negative, and they send the amendment back to step 1 rather than forward to a plan:

MarkerVerdictConsequence
A1 — do custom headers reach every request?ConfirmedR14 stands as written
A4 — is Gw2ForbiddenError.scope usable?Confirmed, 3 of 4 toolsR15/SC18 narrowed
A2 — is a shaped materials result small enough?RefutedR16's default cannot ship
A3 — is the ranking tool affordable?RefutedR13 loses a tool

Per the constitution, "a refuted claim sends the spec back to step 1 rather than being patched." Two claims are refuted, so spec.md needs a revision the human approves — this document does not patch around them. Revision A2 was written on 2026-08-18 and is pending re-approval.

That revision opened one further marker, A5 — does a client-side tool timeout still leave the account reads cached? It now has a verdict below (confirmed, from code).

All five Amendment A markers therefore have verdicts, and the status is complete. The human re-approved Revision A2 on 2026-08-18, which the agent transcribed into spec.md on explicit instruction. The plan gate is open.

The spike for A2/A3 read a real GW2 API key from a file, never printed it, and both the key file and the spike were deleted immediately after the run.

Verified against, and when. All findings dated 2026-08-16. MCP specification revision 2026-07-28 (current, final) at modelcontextprotocol.io. @modelcontextprotocol/sdk1.30.0 (latest published on npm at that date). Claude Code documentation at code.claude.com/docs, and the client itself tested at version 2.1.232. Repository state: branch 017-mcp-server at the commit that added spec.md. Node v26.5.0 for the spikes; apps/api on NestJS 11.1.28, @nestjs/platform-fastify 11.1.28, fastify 5.10.0.

Two discovery spikes were run under the constitution's second exception — a Nest+Fastify MCP server (V2) and a GW2 payload measurement (V3). Both lived in the session scratchpad, outside the repository, and are discarded. Their outputs are recorded below.


Question. spec.md R2 assumes Streamable HTTP is current, HTTP+SSE is deprecated, and that running without sessions is a legitimate choice among alternatives.

Verdict. Confirmed, and stronger than assumed — but against a revision the spec was not written for. Streamable HTTP is current and HTTP+SSE is deprecated. Statelessness is not a choice: revision 2026-07-28 removed protocol-level sessions entirely, so the stateless POST-only shape R2 picked is now the only shape the current revision defines. The spec's decision is vindicated; the spec's wording targets a superseded revision and needs a pass (see Refuted claims).

Evidence. Verified twice — by the discovery agent and then independently by direct fetch of the specification, because this finding rewrites requirements.

  • Current revision is 2026-07-28, final. https://modelcontextprotocol.io/specification/latest resolves to it and cites the normative schema at schema/2026-07-28/schema.ts.
  • https://modelcontextprotocol.io/specification/2026-07-28/basic/transports/streamable-http: "Revision 2026-07-28 changed the behavior of Streamable HTTP… Changes included: Removal of the GET stream endpoint. Removal of protocol-level sessions."
  • Same page, on HTTP+SSE: "Deprecated: The HTTP+SSE transport from protocol version 2024-11-05 has been deprecated since protocol version 2025-03-26… New implementations SHOULD NOT adopt it."
  • Same page, on legacy traffic to a modern-only server: "HTTP GET or DELETE to the MCP endpoint: respond with 405 Method Not Allowed. An Mcp-Session-Id header on a request: ignore it, and do not mint or echo session IDs."
  • Changelog: "Make MCP stateless: remove the initialize/notifications/initialized handshake. Every request now carries its protocol version and client capabilities in _meta."

What else the revision mandates — all confirmed by direct fetch, all new obligations the spec does not currently carry:

ObligationDetail
Required headersMCP-Protocol-Version on every POST; Mcp-Method on all requests; Mcp-Name on tools/call. "These headers are REQUIRED for compliance."
Header/body validationServer MUST reject mismatches with 400 + JSON-RPC -32020 (HeaderMismatch). A missing required header is itself a failure.
Origin validation"Servers MUST validate the Origin header on all incoming connections to prevent DNS rebinding attacks… MUST respond with HTTP 403 Forbidden."
Unknown method404 Not Found + JSON-RPC -32601 — not a 200 carrying an error.
Notifications202 Accepted with no body.
server/discoverMandatory; servers must advertise supported versions, capabilities and identity.
List resultsttlMs and cacheScope required on tools/list.
No resumabilityLast-Event-ID unsupported.

Caveat. Confirmed on paper only. Whether we can build to this revision is V2's problem, and the answer there is no.


V2 — Does the SDK's Streamable HTTP transport work behind Nest's Fastify adapter, in stateless mode? ​

Question. spec.md R2 flags the risk that the SDK is written against Node/Express-shaped req/res while apps/api runs Fastify, which parses the body before a handler sees it.

Verdict. Confirmed — the Fastify handoff is a non-problem. A stateless, POST-only MCP endpoint inside NestJS-on-Fastify works, proven end to end with the real SDK client. The integration risk R2 called out is three lines of glue, not a redesign.

Evidence. Spike: real Nest 11 + @nestjs/platform-fastify 11.1.28 + fastify 5.10.0 — the same versions as apps/api — driven by the SDK's own Client + StreamableHTTPClientTransport against http://127.0.0.1:3999/api/mcp, with app.setGlobalPrefix('api') active.

  • Working handler: new McpServer and new StreamableHTTPServerTransport({ sessionIdGenerator: undefined, enableJsonResponse: true }) per request, then await transport.handleRequest(req.raw, reply.raw, req.body).
  • The one gotcha, and the SDK's answer to it. handleRequest(req, res, parsedBody?) reads the body from the stream only when parsedBody is omitted (dist/esm/server/webStandardStreamableHttp.js:484-487). Omitting it under Fastify yields 400 {"code":-32700,"message":"Parse error: Invalid JSON"} because Fastify already drained the stream. Passing req.body fixes it. No raw-body plugin, no content-type parser removal.
  • A stateless transport serves exactly one request — webStandardStreamableHttp.js:172-176 throws "Stateless transport cannot be reused across requests." Verified: a reused singleton gives 200 then a bodyless 500. A fresh server and transport per request is mandatory, not stylistic.
  • reply.hijack() is not required. Nest's non-passthrough @Res() means Nest never calls reply.send(), and reply.sent reflects raw.writableEnded, which the SDK has already set. Measured after the handler: sent=true raw.writableEnded=true. No double-send.
  • Global interceptors run after the response is on the wire (headersSent=true), so response-shaping interceptors and filters do not apply to this route. They do not break it either.
  • InMemoryTransport.createLinkedPair() (@modelcontextprotocol/sdk/inMemory.js) works with no network listener — good for the tool-layer tests SC1/SC2 want. It does not exercise the HTTP transport, so the Fastify handoff needs its own HTTP-level test; Nest's app.inject() covers it.

Caveat — and it is the important one. The spike negotiated protocolVersion: "2025-06-18" and ran an initialize handshake, which the current revision removed. This is not a spike defect; see F1.

Secondary caveat. The SDK carries hard (non-optional) dependencies on express, hono, @hono/node-server, cors, jose, ajv, eventsource and express-rate-limit — real install weight for a package we use two classes from. Notably it pulls in Express, which this project deliberately does not use.

Behavioural differences from an Express setup, both caused by Fastify parsing first: malformed JSON returns Fastify's 400 {"statusCode":400,"message":"Body is not valid JSON…"} rather than JSON-RPC -32700, and Fastify's 1 MB body limit applies before the transport sees anything.


V3 — How much does field-shaping actually reduce payload? ​

Question. R5 asserts trimmed tool results save meaningful context. An unbacked number is a guess wearing a criterion's clothes.

Verdict. Confirmed for items and prices, weak for recipes. Measured live against api.guildwars2.com.

Evidence. The upstream API returns pretty-printed JSON, so "raw" has two honest readings. Whitespace alone is 39% of the items payload, so the apples-to-apples figure is against minified raw — what a parse-and-re-emit proxy produces. Shaped is measured compact.

Endpointrecordsraw prettyraw minshapedreduction vs minvs pretty
/v2/items5046,217 B28,137 B7,623 B72.9%83.5%
/v2/recipes5027,829 B16,617 B11,415 B31.3%59.0%
/v2/commerce/prices5710,962 B6,857 B1,872 B72.7%82.9%

Approximate tokens (chars/4 estimate, no tokenizer run): items ~7,034 → ~1,906; recipes ~4,154 → ~2,854; prices ~1,714 → ~468.

What dominates raw items — not what R5 claims. Per-field bytes across 50 records: details 9,067 (32.3%), icon 4,778 (17.0%), game_types 2,434 (8.7%), flags 2,388 (8.5%), description 1,615 (5.8%), chat_link 1,300 (4.6%). R5 names description and chat_link as the villains; they are together only 10.4%. The real weight is details, icon and game_types.

A cost we are choosing to pay. flags is 2,388 B of the 7,623 B shaped result — 31% of what we keep, the most expensive retained field. It stays because buyable-vs-gated classification depends on it (docs/project-brief.md, and the classification field in recipe-graph.schema.ts). Dropping it would take items to ~5.2 KB (~81%).

Recipes are the weak case, stated plainly. 31.3%, because the shape keeps ingredients (37.7% of raw) and disciplines — which are most of a recipe record. What it strips is chat_link, time_to_craft_ms, guild_ingredients (almost always []), type and flags. Real, but modest.

Prices beat the guess. 72.7%, from flattening the nested buys/sells objects and dropping whitelisted. Small in absolute terms — ~120 B → ~33 B per record.

Caveat. HTTP gzips the wire, so this is not a bandwidth saving. It is a context-window saving, and the token estimates are the number that matters.

Reproduction. Item ids used (50): the 21 gen-1 legendaries from apps/api/src/legendaries/legendaries.data.ts plus named mats and material-category ids. Recipe ids from /v2/recipes/search?input=19721. Prices: 58 requested, 57 returned (206) — 20796 is not tradeable.


V4 — How does Claude Code connect to a remote streamable-HTTP server with a static bearer token? ​

Question. R10 assumes a documented client configuration exists that carries a bearer header without committing the secret.

Verdict. Confirmed, with two constraints that change requirements. The configuration shape exists and is documented; a plain bearer token is workable, but only if the server is careful about how it refuses.

Evidence.

  • Config shape is type / url / headers. CLI route: claude mcp add --transport http <name> <url> --header "Authorization: Bearer <token>" (code.claude.com/docs/en/mcp). streamable-http is an alias for http. A JSON entry with a url but no type is a configuration error — it is read as a stdio server.
  • claude mcp add defaults to local scope, writing to ~/.claude.json, not the repository. That alone avoids committing a secret if no shared config is needed.
  • Path is unconstrained: the spec requires only "a single HTTP endpoint path… that supports POST", so /api/mcp is fine.

Constraint 1 — the token only stays a token if the server never mentions OAuth. Documented behaviour is what we want: "If you configured headers.Authorization for the server and the server rejects that header, Claude Code reports the connection as failed instead of falling back to OAuth." But when a server also advertises OAuth, the client has been observed pursuing discovery and ignoring the configured header — issue #59467 (closed as duplicate) and issue #80785 (open as of 2026-07-24), where OAuth metadata parsing hard-fails rather than falling back. Consequence for R3: on a bad or missing token return a bare 401 with no WWW-Authenticate header, and serve no/.well-known/oauth-protected-resource or /.well-known/oauth-authorization-server.

Constraint 2 — do not advertise the subscriptions capability. subscriptions/listen is a long-lived POST stream. Claude Code changelog v2.1.233 records a fix for "MCP v2 connections endlessly reopening the subscriptions/listen stream against servers that terminate long-held streams on a fixed timeout (e.g. serverless hosts)". Render's free plan is exactly such a host. Deferred, not current:subscriptions/listen is a 2026-07-28 mechanism, and neither SDK 1.30.0 nor the tested client implements it (F1). This becomes live at the SDK upgrade, and belongs in the migration note rather than in 017's requirements.

Correction to this entry. It concluded from changelog wording ("MCP v2 connections", subscriptions/listen) that Claude Code had moved to 2026-07-28. Direct wire evidence refutes that — see F1. Client 2.1.232 negotiates 2025-11-25 and sends the initialize handshake. Inference from a changelog was treated as evidence; the wire log is the evidence.

Open sub-question, carried forward. ${VAR} expansion inside headers is documented — headers is explicitly listed as an expansion location, with "Authorization": "Bearer ${API_KEY}" as the worked example — but issues #6204 and #51581 report it not substituting for HTTP transport, and no changelog entry closes the loop. Failure mode is nasty: an unset variable still loads and the literal ${VAR} string is sent, producing a confusing 401 rather than a clear error. Documented alternative is headersHelper, which writes a JSON object of headers to stdout, and which also re-runs and retries once on 401/403. This needs an empirical test before R10's documented config depends on it.


V5 — Is a name→id index scoped to what we already hydrate cheap enough? ​

Question. spec.md puts item-name search out of scope on the grounds that an index costs ~350 batched requests. The lead to check was a much narrower index over items this project already hydrates.

Verdict. The 350-request premise is refuted for the scoped index — it costs 2 requests. But the obstacle is not cost, and it is not cache eviction either; it is that nothing persists and nothing warms.

Evidence.

  • Size: 263 distinct ids. 198 legendary roots (apps/api/src/legendaries/legendaries.data.ts:22-66) plus 86 ids reachable by BFS over the committed packages/legendary-recipes/data/gen1-weapons.json (27 recipes, no orphans), overlapping by the 21 gen-1 roots: 198 + 86 − 21 = 263. At the 199-id cap (gw2-client.ts:27) that is 2 batched /v2/items requests — 0.5% of the rejected sweep.
  • The motivating example is inside the set: Mystic Clover is 19675.
  • Eviction is a non-issue. BoundedCache is a size-capped Map with FIFO eviction (bounded-cache.ts:16-53) at CACHE_MAX_ENTRIES = 10_000 (gw2-client.ts:31). ~300 entries is 3% full.
  • Names are already cached. staticCache holds whole parsed entities including name (gw2-client.ts:236, schema gw2.schemas.ts:8-27) with no TTL. ItemDataService holds no state of its own — it is a pass-through (item-data.service.ts:26-36).
  • Nothing warms. No OnModuleInit or onApplicationBootstrap anywhere in apps/api. recipe-graph.service.warm-cache.test.ts:77-92 is misleadingly named — it asserts a second resolve() issues zero fetches; it is not a warming path.
  • Nothing persists. No DB driver, no ORM, no file cache in apps/api/package.json; render.yaml provisions no database and no disk. docs/architecture/stack.md:8 names Postgres as intended but unimplemented.
  • The deploy empties it. Render's free plan spins down on inactivity with a cold start "on the order of a minute" (docs/architecture/deploy.md:84-89), clearing all three caches.
  • No committed names exist. gen1-weapons.json is ids only. This is deliberate: specs/012-app-shell/research.md:189-191 — "none of the scraped names need transcribing into the codebase."

Options, ranked by cost.

  1. Boot-time hydration of the 263 committed ids — 2 batched requests, reuses Gw2Service.items and its no-TTL cache. Needs a boot hook (none exists) and a decision about re-warming after spin-down.
  2. Read whatever staticCache holds — 0 requests, but BoundedCache exposes no iteration (bounded-cache.ts:24-52) so it needs a code change anyway, and after a cold start it holds nothing, so the first agent question misses.
  3. Commit a generated name→id JSON — 0 runtime requests, survives restarts, no persistence layer. Needs a regeneration story for game patches, and contradicts the standing 012 decision.

Second obstacle, smaller. Depth is gen-1 only: 145 of the 198 legendaries are single-node trees until a curation spec deepens them (specs/007-recipe-graph-resolver/research.md:78-92).

This is a scope decision for the human, not a research conclusion. The stated reason for exclusion no longer holds. Whether the tool joins 017 or becomes its own spec is not the agent's call.


F1 — The published SDK does not implement the current spec revision ​

Finding. The MCP specification is at 2026-07-28, but @modelcontextprotocol/sdk 1.30.0 — the latest published version — declares:

LATEST_PROTOCOL_VERSION = '2025-11-25'
DEFAULT_NEGOTIATED_PROTOCOL_VERSION = '2025-03-26'
SUPPORTED_PROTOCOL_VERSIONS = ['2025-11-25','2025-06-18','2025-03-26','2024-11-05','2024-10-07']

2026-07-28 is absent. Confirmed by reading dist/esm/types.js in the installed package and by npm view @modelcontextprotocol/sdk versions (1.30.0 is genuinely latest, not a stale local install). This is corroborated behaviourally: the V2 spike negotiated 2025-06-18 and performed an initialize handshake — a handshake the current revision removed.

This refutes a claim relayed during discovery. The 2026-07-28 release announcement states "All four Tier 1 SDKs speak 2026-07-28 as of today: TypeScript, Python, Go, and C#." The npm artifact contradicts it. The announcement was taken as evidence; the artifact is the evidence.

Why it matters. It splits the feature's foundation:

  • Building on the SDK (an Assumption in spec.md — hand-rolling the protocol is rejected) means shipping a server that speaks 2025-11-25 or older, i.e. with sessions, with initialize, with a GET endpoint — the revision V1 confirms is superseded.
  • V4 found evidence that Claude Code has already moved to the new revision ("MCP v2 connections", subscriptions/listen, changelog v2.1.232/233).

Closed — the premise of the worry was wrong. The client has not moved. Tested live: Claude Code 2.1.232 registered against the V2 spike server (claude mcp add --scope local --transport http spike http://127.0.0.1:3999/api/mcp) reports ✔ Connected, and a real tools/call executed end to end through the CLI into the Nest+Fastify handler. Server-side wire log:

POST /api/mcp   initialize   protocolVersion "2025-11-25"   user-agent claude-code/2.1.232 (sdk-cli)   -> 200
POST /api/mcp   notifications/initialized   mcp-protocol-version: 2025-11-25                           -> 202
GET  /api/mcp   accept: text/event-stream                                                              -> 404 (tolerated silently)
POST /api/mcp   tools/list                                                                             -> 200
POST /api/mcp   tools/call                                                                             -> 200
  • It negotiates 2025-11-25 — the SDK's LATEST_PROTOCOL_VERSION. Client and SDK are on the same revision; there is no protocol break to bridge.
  • It POSTs initialize, the classic handshake.
  • MCP-Protocol-Version is sent on every request after initialize. Mcp-Method and Mcp-Name are never sent — no trace of the 2026-07-28 header contract.
  • It issues exactly one GET (a standalone-SSE probe) which 404s harmlessly and does not affect connection health. It never issues DELETE, because stateless means no session id to terminate.
  • The legacy-fallback path was never exercised, because the client never attempts a modern-path request. Zero 400/404/405 on any POST.

The deferred risk, measured rather than guessed. Simulating a future strictly-modern client by hand against SDK 1.30.0:

  • No initialize, with MCP-Protocol-Version: 2026-07-28 and Mcp-Method: tools/list → 400 {"code":-32000,"message":"Bad Request: Unsupported protocol version: 2026-07-28 (supported versions: 2025-11-25, …)"}. That is a well-formed JSON-RPC error body, which is precisely the shape a fallback-capable client is documented not to retry on — so a strictly-modern client fails hard rather than degrading.
  • Still sending initialize with protocolVersion: 2026-07-28 → SDK 1.30.0 negotiates down to 2025-11-25 and succeeds.

So the risk is real, deferred, and contained inside one dependency. The day Claude Code drops the initialize handshake, the remedy is pnpm up @modelcontextprotocol/sdk — the Nest/Fastify glue (fresh server, fresh transport, handleRequest(req.raw, reply.raw, req.body)) is untouched by protocol revisions.

Consequence for the spec. R2 targets what the SDK and client actually speak — 2025-11-25 — and records 2026-07-28 as a known future migration rather than a requirement. The 2026-07-28 obligations catalogued in V1 (mandatory Mcp-Method/Mcp-Name headers, -32020 validation, server/discover, ttlMs/cacheScope, 202 on notifications, 404+-32601) are not requirements for 017; they are the migration checklist.

Hygiene note. The test touched only ~/.claude.json under the scratchpad project key, and was undone with claude mcp remove spike --scope local; global and project mcpServers verified empty afterwards. Nothing was written into the repository.


A1 — Does an MCP client forward a custom header on every request, or only on initialize? ​

Question. spec.md R14 [A] carries the whole account-tool design on one assumption: that a header set in client configuration (X-GW2-Key) reaches every tools/call, not just the handshake. If clients only send configured headers on initialize, R14 collapses and the credential has to travel some other way. Research V4 already found ${VAR} expansion in the same headers block unreliable, so nothing about this block is assumed twice.

Verdict: CONFIRMED. Every request carries the header, tools/call included.

Method. A throwaway zero-dependency Node HTTP server speaking raw JSON-RPC (initialize, tools/list, tools/call), logging the headers of every request it received. Registered with claude mcp add --scope local --transport http … --header "X-GW2-Key: PROBE-SECRET-123", then driven by a real Claude Code client (claude -p, headless) instructed to call the tool. Ran in the session scratchpad, outside the repository, and is discarded; the local registration was removed afterwards (claude mcp remove) and claude mcp list confirmed empty.

Wire log, Claude Code 2.1.234, 2026-08-18 — one full connect-and-call cycle:

text
POST server/discover           key=YES  Mcp-Protocol-Version: 2026-07-28  Mcp-Method: server/discover
POST initialize                key=YES  (initialize requested protocolVersion 2025-11-25)
POST notifications/initialized key=YES  Mcp-Protocol-Version: 2025-11-25
GET  (standalone SSE probe)    key=YES  Mcp-Protocol-Version: 2025-11-25
POST tools/list                key=YES  Mcp-Protocol-Version: 2025-11-25
POST tools/call                key=YES  Mcp-Protocol-Version: 2025-11-25

6 of 6 requests carried the header. An earlier run of the same probe across two connect cycles gave 11 of 11. The tools/call header set was accept, accept-encoding, connection, content-length, content-type, host, mcp-protocol-version, user-agent, x-gw2-key — the custom header arrives untouched, and the tool's echoed result confirmed the value reached the handler: pong; x-gw2-key=PROBE-SECRET-123.

Consequence for the spec. R14 stands as written. The header transport is viable, and no account tool needs a key parameter.

A caveat this evidence does not cover. One client at one version was tested. The MCP specification does not oblige a client to forward arbitrary configured headers, so this is an observed behaviour of Claude Code 2.1.234, not a guarantee across clients. docs/architecture/mcp.md should say so rather than presenting it as a protocol property.


A1-b — Claude Code 2.1.234 does send Mcp-Method, partially refuting F1 ​

Not a marker — an incidental finding from A1's wire log, recorded because it contradicts a documented claim.

docs/architecture/mcp.md (line 59) records, from F1 against Claude Code 2.1.232/2.1.233, that the client "never sends Mcp-Method or Mcp-Name — no trace of the newer header contract."

That is no longer true. Version 2.1.234 opens every connection with a server/discover POST carrying both Mcp-Protocol-Version: 2026-07-28 and Mcp-Method: server/discover — the 2026-07-28 header contract that R10's migration checklist describes. Only when that probe fails does it fall back to initialize at 2025-11-25.

Scope of the refutation, stated narrowly. Mcp-Method appears only on the server/discover probe. After the fallback, initialize, tools/list and tools/call carry no Mcp-Method and no Mcp-Name. So R2's decision to target 2025-11-25 is unaffected — the client still negotiates and operates on that revision. What changed is that the migration checklist now has a live trigger: the newer contract has begun appearing in a shipping client, one method at a time.

Our server was never tested against this, so it was tested now. SC10's manual record is dated 2026-08-16 against client 2.1.233, which did not send server/discover. Against the real API (MCP_AUTH_TOKEN set, run locally from dist/), the probe is answered:

text
HTTP 400
{"jsonrpc":"2.0","error":{"code":-32000,
 "message":"Bad Request: Unsupported protocol version: 2026-07-28
            (supported versions: 2025-11-25, 2025-06-18, 2025-03-26, 2024-11-05, 2024-10-07)"},
 "id":null}

The client falls back cleanly: claude mcp list → ✔ Connected, and a live tools/call on gw2_items {"ids":[19721]} returned [{"id":19721,"name":"Glob of Ectoplasm","rarity":"Exotic","type":"Trophy","flags":[],"vendor_value":256}] from the live GW2 API. No regression — but the 400/-32000 answer is the SDK's, not ours, and it is what a future client version may stop tolerating. That is the migration checklist's trigger, and it is now closer than F1 implied.


A4 — Does Gw2ForbiddenError.scope carry a usable scope name? ​

Question. spec.md R15 [A] promises that a key missing a scope produces an error naming the scope. That is only implementable if Gw2ForbiddenError.scope reliably holds a real scope name for the endpoints Amendment A exposes. This project has already been wrong once about GW2's auth-failure semantics (docs/gaps/gw2-auth-status.md), so it is checked rather than assumed.

Verdict: CONFIRMED for three of the four account tools; REFUTED for priory_account. R15 needs narrowing — see Consequence.

Answered from the code, no live probe needed: the mapping is deterministic and already covered by tests.

The mechanism works. scopeFromBody (gw2-client.ts:78) parses GW2's documented 403 body { "text": "requires scope <scope>" } with /requires scope (\w+)/, returning null on an absent or unparseable body. authedCachedRead (gw2-client.ts:465-471) maps 401 → Gw2UnauthorizedError and 403 → Gw2ForbiddenError(parsed ?? fallbackScope). The two statuses are kept distinct, which is exactly what the 018 correction in docs/gaps/gw2-auth-status.md established as possible.

Every authenticated read supplies a sensible fallback, so scope is never empty even when the body is missing:

ReadFallback scope
/account/materialsinventories
/account/walletwallet
/account/bankinventories
/account/inventoryinventories
/characterscharacters
characterInventoryinventories

Both branches are already under test: gw2-client.inventory.test.ts:106 asserts the parsed scope (characters), and :119 asserts the fallback when the body is unparseable. R15 therefore needs no new client work — only the mapping in toToolErrorText.

The exception. Gw2Client.account() (gw2-client.ts:190-193) maps both 401 and 403 to Gw2UnauthorizedError, and never constructs a Gw2ForbiddenError:

ts
if (res.status === 401 || res.status === 403) {
  throw new Gw2UnauthorizedError('GW2 rejected the API key');
}

So priory_account cannot name a missing scope. In practice this costs nothing: /v2/account requires the account scope, which is mandatory on every GW2 key and cannot be absent, so a 403 there is not a reachable state. The collapse is harmless — but R15 states the scope-naming rule without qualification, and an unqualified rule that one tool cannot satisfy is a spec that lies.

Worth recording: the collapse is a residue of spec 016's belief that GW2 answers 403 for a bad key. The docstring above it still carries that reasoning ("A 401/403 maps to Gw2UnauthorizedError"), which docs/gaps/gw2-auth-status.md has since corrected — live GW2 returns 401 for a rejected key and 403 only for a missing scope. This is a documentation residue, not a bug: the behaviour is still correct for this endpoint, for the reason given above. Amendment A does not touch it. Noted so the next reader does not mistake it for the pattern to copy — authedCachedRead is.

Consequence for the spec. R15's third row is narrowed to the three tools that can honour it (priory_account_materials, priory_account_wallet, priory_legendary_ranking — the last reaching inventories and characters through getOwnedItems). priory_account maps a rejected key to the invalid-key message and nothing more. No client-side change is needed for any of them.


A2 — How big is a shaped priory_account_materials result? ​

Question. spec.md R16 [A] assumes that dropping count: 0 rows and icon makes an owned-materials result small enough to be a tool response. If it does not, the default needs a bound rather than a filter.

Verdict: REFUTED. R16 is insufficient as written.

Method. Spike against the live GW2 API with a real key on an established account, enriching exactly as AccountService.getMaterials does (items + /materials?ids=all + /commerce/prices), then applying R16's shaping. Measured 2026-08-18. The key was read from a file, never printed, and both the file and the spike are discarded. Cost: 8 upstream requests for one uncached materials read.

rowspayloadrough tokens
Current REST response (all rows, with icon)680114.4 KB~29,300
R16 shaped (count > 0, no icon)46870.6 KB~18,100

Reduction: 38.3% — and that is the problem. 69% of rows are owned, so dropping zeroes removes less than a third of them. The assumption behind R16 was that unowned slots dominate; on a real account they do not.

~18,000 tokens is not a tool result, it is a context transplant. It is roughly 2.4× the entire measured saving of the original four tools (research V3) spent on a single call, and it would evict most of what an agent was holding. R16's default — return every owned row — cannot ship.

What the numbers support instead. The two questions an agent actually asks have very different shapes, and only one of them needs rows:

  • "Do I have enough of X, Y, Z?" → the bounded ids filter R16 already specifies. Cheap, precise, already correct. Keep unchanged.
  • "What is in my storage?" → 468 rows is not the answer to this. A per-category roll-up is: ~9 category rows with counts and total value, which is what /materials?ids=all already provides the categories for, at roughly 2% of the tokens.

This is a spec decision, not a research one, so it is recorded here and referred back to the human rather than resolved: R16's no-argument default must change from "all owned rows" to something bounded, and the aggregate shape is the option the evidence points at.


A3 — What does priory_legendary_ranking cost? ​

Question. spec.md R13 [A] exposes RankingService.rank as one tool call. spec.md stated the consequence in advance: "If it is too slow or too costly, it does not ship as a tool in this amendment."

Verdict: REFUTED. It does not ship as specified.

Method. Same spike, same account, timing the getOwnedItems fan-out (/characters, one /characters/:name/inventory per character, /account/bank, /account/inventory) against the live API with a warm connection.

MeasureValue
Characters on the account19
Upstream requests per call22 (1 + 19 + 2)
Elapsed31,443 ms
Share of the 300/min budget7.3% per call

31 seconds, and that is the floor, not the figure. This timed getOwnedItems only. RankingService.rank additionally resolves recipe graphs and prices across the legendary set on top of it, so the real tool is slower than what was measured. The honest number is "over 31 s".

Why that is disqualifying rather than merely slow. A tool call that takes half a minute exceeds typical MCP client tool timeouts, so the common outcome is not a slow answer but a failed one. On Render's free tier the instance also spins down, so a first call pays cold start on top. And at 7.3% of the shared minute-budget per invocation, an agent that retries a timeout twice has spent a fifth of the budget every other consumer of the token depends on.

Scaling note: the request count is 1 + characters + 2. This account has 19 characters; the cost is linear in that number and is not a property of the code we could tune away. The ACCOUNT_TTL_MS cache makes a second call within the window cheap, but the first call in every window pays full price, and that is the call an agent makes.

What the evidence supports. Dropping priory_legendary_ranking from Amendment A — which spec.md pre-authorised — and leaving the ranking to the web UI, where a 30-second load behind a spinner is a different proposition from a blocking tool call. The remaining five tools are unaffected: A2's materials read costs 8 requests, and the other four are single reads.

Recorded rather than decided: removing a tool changes R13, SC14's tool count and the traceability rows, and that is the human's call.


A5 — Does a client-side timeout still leave the account reads cached? ​

Question. spec.md R18 [A] advises one retry after a priory_legendary_ranking timeout, on the theory that the first call's work completes server-side and warms the 5-minute account cache. Revision A2 recorded that as inferred, not measured. This closes it.

Verdict: CONFIRMED, from code, with one environmental caveat. Four independent links, each checkable:

  1. No cancellation channel exists. grep for AbortSignal, AbortController and signal: across apps/api/src (excluding tests) returns nothing. fetchWithRetry (gw2-client.ts:485) forwards only init — headers — to this.fetchFn. There is no way for a disconnect to reach the outbound request, because no such wiring was ever built.
  2. The sockets are unrelated. The client→us connection and our→api.guildwars2.com connection are separate TCP connections. Closing the first cannot affect the second; nothing in between is watching.
  3. The handler is not cancellable. mcp.controller.ts's reply.raw.on('close', …) closes the transport and the MCP server. It does not — and in JavaScript cannot — cancel the in-flight handler promise. The tool handler runs to completion; only its result has nowhere to go.
  4. The cache write does not depend on the response. authedCachedRead calls this.accountCache.set(key, parsed, ACCOUNT_TTL_MS) immediately after parseOrThrow and before returning to the caller. The write happens whether or not anything ever reads the return value.

Together: a timed-out first call still populates the cache for ACCOUNT_TTL_MS (5 minutes), so R18's retry advice holds. No live measurement was needed — every link is a property of code in the repository, and a spike could only have confirmed what the four already determine.

The caveat, which R18's description text must respect. accountCache is in-process memory. On Render's free tier the instance spins down when idle, and any restart empties it. So the retry advice is true within a warm instance and false across a spin-down — where the retry pays the full >31 s again, plus cold start. The description should therefore say a retry is likely fast, not that it is fast.

Consequence for the spec. R18 stands; its [NEEDS VERIFICATION: A5] marker is discharged. The wording constraint above is a requirement on the description text, not a new design question.


F2 — Two GW2 ids were wrong in working notes ​

Finding. While briefing the V3 measurement, two ids were asserted from memory and both were wrong: 20796 is Philosopher's Stone, not Mystic Clover (which is 19675), and 30703 is Sunrise, not The Bifrost. Caught by the agent checking names against /v2/items rather than accepting the brief.

Why it matters. Nothing in spec.md asserts either name, so no requirement changes. It is recorded because it is direct evidence for the problem V5 addresses: the id↔name mapping is exactly what is routinely got wrong from memory, which is the case for the name-lookup tool.

F3 — /v2/recipes cannot describe a legendary ​

Finding. /v2/recipes/search?output=30703 returns []. Mystic Forge recipes are absent from the API, as docs/project-brief.md and docs/architecture/gw2-api.md already record — re-confirmed live.

Why it matters. The gw2_recipes and gw2_recipe_search tool descriptions (R4) must say so. An agent that asks the API for a legendary's recipe gets an empty array, not an error, and an empty array reads as "no such recipe" rather than "this API cannot answer that". The curated table is the only source, and the tool descriptions are where a model would learn it.


F4 — A repository claim in spec.md was made against the wrong checkout ​

Finding. spec.md's original out-of-scope entry asserted that account support did not exist and that 016-account-api-key was an unmerged branch. Both are false. origin/main is at 3a77065 — "016 — GW2 API key: connect & validate (#18)", and this worktree branched from it, so apps/api/src/account/ is present: GET /api/account takes a player key via Authorization: Bearer <key> and validates it against /v2/account, and Gw2Service.account(apiKey) exists (gw2.service.ts).

How it happened. The grep that produced "zero account/auth code in apps/api" was run in the shared checkout at 071462b — a stale local main — before EnterWorktree created this worktree from origin/main. The tree was never re-checked after the move.

What survives. AccountService is pure pass-through and stores nothing, so V5's "nothing is persisted anywhere in this repo" stands, and the plan is structurally unaffected. The other repository claims in V5 and V2 were made by agents reading this worktree, so they are against the correct base.

Why it matters. The out-of-scope entry has been corrected: account tools stay out of 017 not because intake is missing, but because the player's GW2 key and this server's bearer token would both want the Authorization header, so the key needs another route — design work, not a thin wrapper.

The general lesson, worth more than the fix. A repository claim is only as good as the tree it was read from. After entering a worktree, re-establish any fact gathered before the move.

Refuted claims ​

An intermediate conclusion drawn during this research was itself wrong, and is withdrawn. After V1 landed, the working conclusion was that spec.md was written against a superseded protocol and had to return to step 1 — because P1, P2 and SC1 describe a client that "completes initialization", and 2026-07-28 removed the initialize handshake. The interop evidence in F1 dissolves that: neither the SDK we build on nor the client that consumes it implements 2026-07-28, both speak 2025-11-25, and initialize is exactly what happens on the wire. The spec's initialize wording is correct for the protocol 017 will actually implement. Recorded because the reasoning was published before the evidence arrived, and because it is the same error as V4's: an authoritative document was taken as evidence of deployed behaviour.

No spec claim was refuted outright. Three need correction, and one decision must be added — targeted edits, not a rewrite:

spec.mdCorrectionSource
R2Must state the targeted protocol revision (2025-11-25, what the SDK and client speak) and record 2026-07-28 as a known future migration. Statelessness gains a second, stronger justification: a stateless SDK transport cannot be reused across requests, so a fresh server and transport per request is mandatory. 405 on GET/DELETE stays, now as tidiness — the client's one GET currently 404s harmlessly.V1, V2, F1
R3Add: on a bad or missing token return a bare 401 — no WWW-Authenticate header, and no .well-known/oauth-* endpoints — or the client may pursue OAuth discovery and ignore the configured header.V4
R5Rewrite the field attribution. It names description and chat_link; measured, those are 10.4% combined, while details + icon + game_types are 58%. Cite 72.9% (vs minified), not the 83.5% pretty-print figure, and frame the win as context-window, not bandwidth.V3
R10The ${VAR}-in-headers mechanism is documented but has a live history of not substituting. Either verify empirically during implementation or use headersHelper / local scope.V4

Requirements that do NOT apply to 017. The 2026-07-28 obligations catalogued in V1 — mandatory Mcp-Method / Mcp-Name headers, -32020 header/body validation, server/discover, ttlMs / cacheScope, 202 on notifications, 404 + -32601, the subscriptions reconnect hazard — are not requirements now. They are the SDK-upgrade migration checklist, and belong in docs/architecture/mcp.md as such. Origin validation is the exception worth adopting early: it is cheap, it is a genuine DNS-rebinding defence, and apps/api already has a CORS allow-list to build on.

The out-of-scope entry for item-name search is refuted in its reasoning (V5): the 350-request figure does not apply to the scoped index, which costs 2 requests. The exclusion may still stand on other grounds — the boot-hook and cold-start questions are real. That is a scope decision for the human, not a research conclusion.

Graduation ​

Candidates to move to docs/architecture/ at step 6, so the next spec does not re-derive them:

  • docs/architecture/mcp.md (new, per R10): the protocol-revision-versus-SDK split (F1), the stateless-transport-per-request constraint (V2), the Fastify req.body handoff (V2), the no-OAuth-hints rule for bearer tokens (V4), and the do-not-advertise-subscriptions rule (V4).
  • docs/architecture/gw2-api.md: the measured payload composition (V3) — details, icon and game_types dominate — and the re-confirmed fact that /v2/recipes/search?output= is empty for legendaries (F3).
  • docs/gaps/: nothing warms and nothing persists, and the free-tier spin-down empties every cache (V5). This is cross-cutting, affects any future caching or index work, and is not 017's to fix.