Skip to content

Report Copilot's prompt-too-long error the way Claude Code recognizes it - #166

Merged
D0n9X1n merged 4 commits into
mainfrom
fix/prompt-too-long
Oct 3, 2026
Merged

D0n9X1n merged 4 commits into
mainfrom
fix/prompt-too-long

Conversation

@D0n9X1n

@D0n9X1n D0n9X1n commented Oct 3, 2026 •

Copy link
Copy Markdown
Owner

Summary

Copilot answers a prompt over the model's max_prompt_tokens with HTTP 400:

{"error":{"message":"prompt token count of 131008 exceeds the limit of 128000","code":"model_max_prompt_tokens_exceeded"}}

Claude Code 2.1.288 recognizes an overflow by prompt is too long or input is too long for requested model, and reads the counts with prompt is too long[^0-9]*(\d+)\s*tokens?\s*>\s*(\d+). The relay returned Copilot's body unchanged to a JSON request, and on a stream translateErrorToClaudeErrorEvent replaced it with An unexpected error occurred during streaming.

The relay now reports it as Anthropic's API does, with HTTP 400:

{"type":"error","error":{"type":"invalid_request_error","message":"prompt is too long: 131008 tokens > 128000 maximum"}}
  • toPromptTooLongError in src/copilot/client.ts recognizes the overflow where a failed upstream response becomes an error: logUpstreamError in src/copilot/chat.ts (/chat/completions, /responses) and createNativeMessages in src/copilot/native.ts (/v1/messages). It requires HTTP 400 and a JSON body whose error.code is model_max_prompt_tokens_exceeded, or Anthropic's own invalid_request_error envelope whose message starts with prompt is too long. It reads only the two counts, with anchored patterns, as safe integers.
  • PromptTooLongError in src/lib/error.ts is an HTTPError that carries that body. When the counts cannot be read, the message is prompt is too long.
  • JSON requests: src/routes/claude.ts already returns an HTTPError's body and status, so the client gets HTTP 400 with that body. No route change.
  • Streaming: translateErrorToClaudeErrorEvent in src/claude/stream.ts sends invalid_request_error with the same message as the SSE error event. The route still appends (request_id=<id>), which the count pattern does not read.
  • Every other error keeps its handling, including UpstreamToolInputError's client-safe message. The existing error log lines stay; logUpstreamError now returns the error to throw instead of its detail.
  • Docs, EN and ZH together: the overflow case in wiki/EN-Logging-Troubleshooting.md and wiki/ZH-Logging-Troubleshooting.md, the mapping in wiki/EN-Internals.md and wiki/ZH-Internals.md.

Closes #158

Streaming: the SSE error event, not a held-back HTTP 400

(a), waiting for Copilot's status before opening the stream, is not clean here. The relay opens the SSE stream before the inference call. POST /v1/messages streaming opens before delayed upstream without ping in tests/integration/claude-routes.test.ts requires the SSE response within 500 ms while the mocked upstream withholds its response; its comment gives the reason: "Claude Code can cancel requests after 60s before response headers." Waiting for Copilot's status would hold the relay's response headers until Copilot answers and break that test.

(b) is what Claude Code 2.1.288 acts on. Read-only inspection of the installed binary (~/.local/bin/claude.exe, which embeds VERSION:"2.1.288"). The live run under Validation ran Claude Code end to end.

Its bundled SDK throws an SSE error event as an APIError with no status:

if(x.event==="error"){let T=of(x.data)??x.data,D=T?.error?.type;throw new xt(void 0,T,void 0,n.headers,D)}

With no status and no top-level message, the event's JSON becomes the error message:

static makeMessage(n,e,t){let r=e?.message?typeof e.message==="string"?e.message:JSON.stringify(e.message):e?JSON.stringify(e):t;if(n&&r)return`${n} ${r}`;if(n)return`${n} status code (no body)`;if(r)return r;return"(no status code or body)"}

The overflow check reads only that message, not the status or the error type:

function b6n(e){let n=e.toLowerCase();return n.includes("prompt is too long")||n.includes("input is too long for requested model")}
function NA(e){if(!(e instanceof Error))return!1;return b6n(e.message)||px(e.message,"prompt_too_long")}

The query path's final catch around the attempt unwraps a retry wrapper to originalError and converts the error with f2, which calls the API-error converter xeo:

let Xi=xi,Ea=V.model;if(xi instanceof rc)Xi=xi.originalError,Ea=xi.retryContext.model;
yield{...f2(Xi,Ea,{messages:e,messagesForAPI:sa,requestId:nl}),perTurnEffort:mE,perTurnTiming:FM,...M3e&&{streamWatchdogGaveUp:!0}}
function f2(e,n,r){let s=xeo(e,n,r);

xeo turns it into Claude Code's prompt-too-long error (uL="Prompt is too long") and keeps the message as errorDetails, from which the counts are read:

if(NA(e)||HO(e))return ds({content:uL,error:"invalid_request",errorDetails:e.message})
function sdt(e){let n=e.match(/prompt is too long[^0-9]*(\d+)\s*tokens?\s*>\s*(\d+)/i);return{actualTokens:n?parseInt(n[1],10):void 0,limitTokens:n?parseInt(n[2],10):void 0}}

oLt turns those counts into a token gap, and a last assistant message that is this error (qye) is what the query loop's reactive compaction (trigger:"ptl") handles.

On a stream error, Claude Code either rethrows it (Error streaming (non-streaming fallback disabled)) or resends the turn without streaming (Error streaming, falling back to non-streaming mode). The resend reaches the relay's JSON path and gets the HTTP 400 above; the SDK's 400 error class (if(n===400)return new WEt(n,o,t,r,i)) builds its message the same way, 400 {"type":"error",...}, which the same check matches. In the live run, Claude Code took the resend: one streaming request, then one non-streaming request that got HTTP 400.

The tests check the produced text with Claude Code's own phrase and count checks in each form it reaches the client: the message, the SSE event's message with the appended request id, the event's JSON, and 400 <body>.

Under (b) the stream still opens before the inference call, sends no ping, and aborts as before; only this error's type and message change.

Scope

Claude Code's POST /v1/messages, streaming and JSON, over all three upstream APIs (/chat/completions, /responses, /v1/messages). No route, config or public API change.

Validation

  • npm run typecheck
  • npm run test:unit: 978 tests, 948 pass, 0 fail, 30 skipped (existing tests that skip on Windows; none of them new)
  • npm run test:integration: 406 tests, 406 pass, 0 fail
  • npm run build
  • Live, against real Copilot (below): Copilot rejected every over-limit request with HTTP 400, and every response from 75b109a matched Claude Code's phrase check and its count pattern.

Run locally on Windows 11 with Node v24.19.0, on c4f88d7: the tree rebased onto fa7ffb8 (after #155, #160, #164 and #165; the rebase had no conflicts), plus a test and a docs commit that pin the live evidence.

Live run, against real Copilot on isolated test relays: 127.0.0.1:4255 on main (fa7ffb8) and 127.0.0.1:4256 on 75b109a; gptModel mai-code-1.1-flash, opusModel claude-haiku-4.5, so a request for claude-opus-5.5 routes to claude-haiku-4.5. Prompts: "hello " repeated 230000 times (claude-haiku-4.5) or 140000 times (mai-code-1.1-flash). It was not repeated for the two commits after 75b109a, which change only a test and docs.

Route, model Mode main (fa7ffb8) 75b109a
/chat/completions, claude-haiku-4.5 JSON HTTP 400 text/plain, Copilot's body unchanged: model_max_prompt_tokens_exceeded, prompt token count of 230008 exceeds the limit of 136000 HTTP 400 application/json, invalid_request_error: prompt is too long: 230008 tokens > 136000 maximum
same stream HTTP 200, event: error, api_error: An unexpected error occurred during streaming. (request_id=…) HTTP 200, event: error, invalid_request_error: prompt is too long: 230008 tokens > 136000 maximum (request_id=…)
/responses, mai-code-1.1-flash JSON HTTP 400 text/plain, Copilot's body unchanged: model_max_prompt_tokens_exceeded, prompt token count of 140001 exceeds the limit of 128000 HTTP 400 application/json: prompt is too long: 140001 tokens > 128000 maximum
same stream the generic api_error event invalid_request_error event: prompt is too long: 140001 tokens > 128000 maximum (request_id=…)
/v1/messages, claude-haiku-4.5 JSON HTTP 400 application/json, Copilot's body unchanged: its code in Anthropic's envelope and wording, prompt is too long: 230024 tokens > 200000 maximum, with a top-level request_id HTTP 400 application/json: prompt is too long: 230024 tokens > 200000 maximum
same stream the generic api_error event invalid_request_error event: prompt is too long: 230024 tokens > 200000 maximum (request_id=…)

Claude Code 2.1.288 end to end: claude -p with its own CLAUDE_CONFIG_DIR, ANTHROPIC_BASE_URL at the test relay, model mai-code-1.1-flash, and CLAUDE_CODE_MAX_CONTEXT_TOKENS=1000000 so it sends the turn instead of blocking it locally. Copilot counted the prompt as 168145 tokens.

  • main: one streaming request (generic error event), then three non-streaming requests, each HTTP 400. Result: terminal_reason=api_error, "API Error: 400 prompt token count of 168145 exceeds the limit of 128000".
  • 75b109a: one streaming request (mapped error event), then one non-streaming request, HTTP 400. Result: terminal_reason=prompt_too_long, "Prompt is too long · the request is ~168145 tokens (limit 128000) but this conversation is only ~78395 tokens — the rest is system prompt, tool definitions, and attachment content. A single-exchange conversation cannot be compacted; reduce attached files/tools or start with less context."

New coverage:

  • tests/unit/prompt-too-long.test.ts: recognition, count parsing and message formatting; Claude Code's checks against every form the client receives; unreadable counts; Anthropic's native shape; the body Copilot sent on /v1/messages in the live run, rebuilt without its request_id; upstream errors that are not an overflow (413 and 500, Copilot's wording without its code, another code or type, the inner error without the envelope, non-object bodies); the stream mapping.
  • tests/integration/claude-routes.test.ts, the [Bug]: Claude Code does not recognize Copilot's prompt-too-long error, and streaming replaces it with a generic one #158 tests, with a mocked upstream: Copilot's 400 on /chat/completions and /responses, JSON and streaming; another upstream 400 keeps its JSON body and stays generic on a stream, on the chat and the native route; the native route with Anthropic's and Copilot's overflow shapes. createTestProxy takes config overrides for the native route.

Notes

  • No config, auth or token change. No log line is added or removed. On a stream, the existing Error during Claude stream request: line now records a PromptTooLongError; it carries the counts and the upstream message, never a credential or prompt content.
  • Judgment calls:
    • toPromptTooLongError sits next to toCopilotAbortHTTPError in src/copilot/client.ts; PromptTooLongError is in src/lib/error.ts, so src/claude/stream.ts checks it without importing from src/copilot/.
    • Anthropic's shape is recognized only with the full envelope (type: "error", error.type: "invalid_request_error") and HTTP 400. A near miss stays generic on a stream, as before.
    • The message is rebuilt from the two counts and never forwarded, so a JSON client of a recognized native body no longer gets upstream's request_id body field. Claude Code reads that field as a fallback request id (Xi.requestID||Xi.error?.request_id), and the relay does not forward upstream's request-id header on this path either.
    • On the native route, the live body carried Copilot's code inside Anthropic's envelope. The recognizer also accepts either one alone; the integration tests cover each shape separately.
    • For a native JSON overflow, the request summary's error= now shows the upstream message instead of the first 240 characters of the body. The chat routes keep their existing detail.
    • src/copilot/responses.ts has no failed-response path; its in-stream failures stay a generic 502. The WebSearch bridge's own /responses call is unchanged: a failure there becomes a failed search result, not a relay error. An error event inside a native stream that has already started is not inspected.
  • Verified live (see Validation): Claude Code was run, and real Copilot calls were made; Copilot rejected every over-limit request with HTTP 400. The native overflow body is now observed, and tests/unit/prompt-too-long.test.ts pins it.
  • Still not verified: reactive compaction in a multi-turn conversation (the live run was a single exchange, which Claude Code says it cannot compact), and what the native route does for claude-haiku-4.5 between 136000 and 200000 tokens (not tested, to avoid a real inference).

Checklist

  • Public API remains Claude Code-only unless intentionally changed.
  • Config changes are reflected in config.default.yaml, README, and wiki/ (EN and ZH). (No config change.)
  • Logs do not expose tokens.
  • Integration tests mock upstream Copilot; they do not call real Copilot services.

After merge

  • Remove inactive local/remote feature branches and temporary worktrees, prune stale refs, and report preserved work; follow the Development Wiki cleanup rules.

🤖 Generated with Claude Code

D0n9X1n and others added 2 commits October 3, 2026 09:23
Copilot rejects a prompt over the model's max_prompt_tokens with HTTP 400,
model_max_prompt_tokens_exceeded and "prompt token count of N exceeds the
limit of M". Claude Code recognizes an overflow only by "prompt is too long",
so it did not see this one, and on a stream the relay replaced it with "An
unexpected error occurred during streaming."

toPromptTooLongError recognizes the rejection where a failed upstream response
becomes an error: logUpstreamError for /chat/completions and /responses, and
createNativeMessages for /v1/messages, which can also carry Anthropic's own
invalid_request_error. It reads only the two counts and returns
PromptTooLongError, whose body is what Anthropic's API sends:
invalid_request_error, "prompt is too long: N tokens > M maximum". A JSON
request gets it with HTTP 400; a stream sends the same type and message as its
SSE error event. Every other error keeps its handling.

Closes #158

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Internals names where the overflow is recognized, the body the client gets,
and why a stream reports it as its SSE error event. Logging and
troubleshooting shows what the client and the log record for it, and what to
do about it. The EN and ZH pages change together.

Closes #158

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@D0n9X1n D0n9X1n added this to the v0.4.6 milestone Oct 3, 2026
@D0n9X1n D0n9X1n added bug Something isn't working dev:windows Taken by the Windows development session labels Oct 3, 2026
D0n9X1n and others added 2 commits October 3, 2026 11:00
In a live run, Copilot answered an over-limit prompt on /v1/messages with
its own code inside Anthropic's envelope and wording, with ">" escaped and
a top-level request_id and type. The unit test keeps that body, with a
placeholder request id, and asserts PromptTooLongError with 230024 and
200000 and a rebuilt body that carries no request_id.

Closes #158

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Internals shows the body Copilot sent on /v1/messages for a prompt over the
limit, with a placeholder request id, and that it named a limit of 200000
for claude-haiku-4.5 there while /chat/completions enforced the catalog's
136000. Only the /chat/completions and /responses wording fails Claude
Code's check, and the recognizer accepts either form on any route, so the
pages no longer say the native route carries Anthropic's envelope instead
of Copilot's code. The EN and ZH pages change together.

Closes #158

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@D0n9X1n
D0n9X1n merged commit 2d49599 into main Oct 3, 2026
6 checks passed
@D0n9X1n
D0n9X1n deleted the fix/prompt-too-long branch October 3, 2026 18:19
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working dev:windows Taken by the Windows development session

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: Claude Code does not recognize Copilot's prompt-too-long error, and streaming replaces it with a generic one

1 participant