Repository navigation
Report Copilot's prompt-too-long error the way Claude Code recognizes it - #166
Merged
Merged
Conversation
Copilot rejects a prompt over the model's max_prompt_tokens with HTTP 400, model_max_prompt_tokens_exceeded and "prompt token count of N exceeds the limit of M". Claude Code recognizes an overflow only by "prompt is too long", so it did not see this one, and on a stream the relay replaced it with "An unexpected error occurred during streaming." toPromptTooLongError recognizes the rejection where a failed upstream response becomes an error: logUpstreamError for /chat/completions and /responses, and createNativeMessages for /v1/messages, which can also carry Anthropic's own invalid_request_error. It reads only the two counts and returns PromptTooLongError, whose body is what Anthropic's API sends: invalid_request_error, "prompt is too long: N tokens > M maximum". A JSON request gets it with HTTP 400; a stream sends the same type and message as its SSE error event. Every other error keeps its handling. Closes #158 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Internals names where the overflow is recognized, the body the client gets, and why a stream reports it as its SSE error event. Logging and troubleshooting shows what the client and the log record for it, and what to do about it. The EN and ZH pages change together. Closes #158 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
In a live run, Copilot answered an over-limit prompt on /v1/messages with its own code inside Anthropic's envelope and wording, with ">" escaped and a top-level request_id and type. The unit test keeps that body, with a placeholder request id, and asserts PromptTooLongError with 230024 and 200000 and a rebuilt body that carries no request_id. Closes #158 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Internals shows the body Copilot sent on /v1/messages for a prompt over the limit, with a placeholder request id, and that it named a limit of 200000 for claude-haiku-4.5 there while /chat/completions enforced the catalog's 136000. Only the /chat/completions and /responses wording fails Claude Code's check, and the recognizer accepts either form on any route, so the pages no longer say the native route carries Anthropic's envelope instead of Copilot's code. The EN and ZH pages change together. Closes #158 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This was referenced Oct 3, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Copilot answers a prompt over the model's
max_prompt_tokenswith HTTP 400:{"error":{"message":"prompt token count of 131008 exceeds the limit of 128000","code":"model_max_prompt_tokens_exceeded"}}Claude Code 2.1.288 recognizes an overflow by
prompt is too longorinput is too long for requested model, and reads the counts withprompt is too long[^0-9]*(\d+)\s*tokens?\s*>\s*(\d+). The relay returned Copilot's body unchanged to a JSON request, and on a streamtranslateErrorToClaudeErrorEventreplaced it withAn unexpected error occurred during streaming.The relay now reports it as Anthropic's API does, with HTTP 400:
{"type":"error","error":{"type":"invalid_request_error","message":"prompt is too long: 131008 tokens > 128000 maximum"}}toPromptTooLongErrorinsrc/copilot/client.tsrecognizes the overflow where a failed upstream response becomes an error:logUpstreamErrorinsrc/copilot/chat.ts(/chat/completions,/responses) andcreateNativeMessagesinsrc/copilot/native.ts(/v1/messages). It requires HTTP 400 and a JSON body whoseerror.codeismodel_max_prompt_tokens_exceeded, or Anthropic's owninvalid_request_errorenvelope whose message starts withprompt is too long. It reads only the two counts, with anchored patterns, as safe integers.PromptTooLongErrorinsrc/lib/error.tsis anHTTPErrorthat carries that body. When the counts cannot be read, the message isprompt is too long.src/routes/claude.tsalready returns anHTTPError's body and status, so the client gets HTTP 400 with that body. No route change.translateErrorToClaudeErrorEventinsrc/claude/stream.tssendsinvalid_request_errorwith the same message as the SSEerrorevent. The route still appends(request_id=<id>), which the count pattern does not read.UpstreamToolInputError's client-safe message. The existing error log lines stay;logUpstreamErrornow returns the error to throw instead of its detail.wiki/EN-Logging-Troubleshooting.mdandwiki/ZH-Logging-Troubleshooting.md, the mapping inwiki/EN-Internals.mdandwiki/ZH-Internals.md.Closes #158
Streaming: the SSE
errorevent, not a held-back HTTP 400(a), waiting for Copilot's status before opening the stream, is not clean here. The relay opens the SSE stream before the inference call.
POST /v1/messages streaming opens before delayed upstream without pingintests/integration/claude-routes.test.tsrequires the SSE response within 500 ms while the mocked upstream withholds its response; its comment gives the reason: "Claude Code can cancel requests after 60s before response headers." Waiting for Copilot's status would hold the relay's response headers until Copilot answers and break that test.(b) is what Claude Code 2.1.288 acts on. Read-only inspection of the installed binary (
~/.local/bin/claude.exe, which embedsVERSION:"2.1.288"). The live run under Validation ran Claude Code end to end.Its bundled SDK throws an SSE
errorevent as anAPIErrorwith no status:With no status and no top-level
message, the event's JSON becomes the error message:The overflow check reads only that message, not the status or the error type:
The query path's final catch around the attempt unwraps a retry wrapper to
originalErrorand converts the error withf2, which calls the API-error converterxeo:xeoturns it into Claude Code's prompt-too-long error (uL="Prompt is too long") and keeps the message aserrorDetails, from which the counts are read:oLtturns those counts into a token gap, and a last assistant message that is this error (qye) is what the query loop's reactive compaction (trigger:"ptl") handles.On a stream error, Claude Code either rethrows it (
Error streaming (non-streaming fallback disabled)) or resends the turn without streaming (Error streaming, falling back to non-streaming mode). The resend reaches the relay's JSON path and gets the HTTP 400 above; the SDK's 400 error class (if(n===400)return new WEt(n,o,t,r,i)) builds its message the same way,400 {"type":"error",...}, which the same check matches. In the live run, Claude Code took the resend: one streaming request, then one non-streaming request that got HTTP 400.The tests check the produced text with Claude Code's own phrase and count checks in each form it reaches the client: the message, the SSE event's message with the appended request id, the event's JSON, and
400 <body>.Under (b) the stream still opens before the inference call, sends no ping, and aborts as before; only this error's type and message change.
Scope
Claude Code's
POST /v1/messages, streaming and JSON, over all three upstream APIs (/chat/completions,/responses,/v1/messages). No route, config or public API change.Validation
npm run typechecknpm run test:unit: 978 tests, 948 pass, 0 fail, 30 skipped (existing tests that skip on Windows; none of them new)npm run test:integration: 406 tests, 406 pass, 0 failnpm run build75b109amatched Claude Code's phrase check and its count pattern.Run locally on Windows 11 with Node v24.19.0, on
c4f88d7: the tree rebased ontofa7ffb8(after #155, #160, #164 and #165; the rebase had no conflicts), plus a test and a docs commit that pin the live evidence.Live run, against real Copilot on isolated test relays: 127.0.0.1:4255 on
main(fa7ffb8) and 127.0.0.1:4256 on75b109a; gptModelmai-code-1.1-flash, opusModelclaude-haiku-4.5, so a request forclaude-opus-5.5routes toclaude-haiku-4.5. Prompts: "hello " repeated 230000 times (claude-haiku-4.5) or 140000 times (mai-code-1.1-flash). It was not repeated for the two commits after75b109a, which change only a test and docs.main(fa7ffb8)75b109a/chat/completions,claude-haiku-4.5text/plain, Copilot's body unchanged:model_max_prompt_tokens_exceeded,prompt token count of 230008 exceeds the limit of 136000application/json,invalid_request_error:prompt is too long: 230008 tokens > 136000 maximumevent: error,api_error:An unexpected error occurred during streaming. (request_id=…)event: error,invalid_request_error:prompt is too long: 230008 tokens > 136000 maximum (request_id=…)/responses,mai-code-1.1-flashtext/plain, Copilot's body unchanged:model_max_prompt_tokens_exceeded,prompt token count of 140001 exceeds the limit of 128000application/json:prompt is too long: 140001 tokens > 128000 maximumapi_erroreventinvalid_request_errorevent:prompt is too long: 140001 tokens > 128000 maximum (request_id=…)/v1/messages,claude-haiku-4.5application/json, Copilot's body unchanged: its code in Anthropic's envelope and wording,prompt is too long: 230024 tokens > 200000 maximum, with a top-levelrequest_idapplication/json:prompt is too long: 230024 tokens > 200000 maximumapi_erroreventinvalid_request_errorevent:prompt is too long: 230024 tokens > 200000 maximum (request_id=…)Claude Code 2.1.288 end to end:
claude -pwith its ownCLAUDE_CONFIG_DIR,ANTHROPIC_BASE_URLat the test relay, modelmai-code-1.1-flash, andCLAUDE_CODE_MAX_CONTEXT_TOKENS=1000000so it sends the turn instead of blocking it locally. Copilot counted the prompt as 168145 tokens.main: one streaming request (generic error event), then three non-streaming requests, each HTTP 400. Result:terminal_reason=api_error, "API Error: 400 prompt token count of 168145 exceeds the limit of 128000".75b109a: one streaming request (mapped error event), then one non-streaming request, HTTP 400. Result:terminal_reason=prompt_too_long, "Prompt is too long · the request is ~168145 tokens (limit 128000) but this conversation is only ~78395 tokens — the rest is system prompt, tool definitions, and attachment content. A single-exchange conversation cannot be compacted; reduce attached files/tools or start with less context."New coverage:
tests/unit/prompt-too-long.test.ts: recognition, count parsing and message formatting; Claude Code's checks against every form the client receives; unreadable counts; Anthropic's native shape; the body Copilot sent on/v1/messagesin the live run, rebuilt without itsrequest_id; upstream errors that are not an overflow (413 and 500, Copilot's wording without its code, another code or type, the inner error without the envelope, non-object bodies); the stream mapping.tests/integration/claude-routes.test.ts, the [Bug]: Claude Code does not recognize Copilot's prompt-too-long error, and streaming replaces it with a generic one #158 tests, with a mocked upstream: Copilot's 400 on/chat/completionsand/responses, JSON and streaming; another upstream 400 keeps its JSON body and stays generic on a stream, on the chat and the native route; the native route with Anthropic's and Copilot's overflow shapes.createTestProxytakes config overrides for the native route.Notes
Error during Claude stream request:line now records aPromptTooLongError; it carries the counts and the upstream message, never a credential or prompt content.toPromptTooLongErrorsits next totoCopilotAbortHTTPErrorinsrc/copilot/client.ts;PromptTooLongErroris insrc/lib/error.ts, sosrc/claude/stream.tschecks it without importing fromsrc/copilot/.type: "error",error.type: "invalid_request_error") and HTTP 400. A near miss stays generic on a stream, as before.request_idbody field. Claude Code reads that field as a fallback request id (Xi.requestID||Xi.error?.request_id), and the relay does not forward upstream'srequest-idheader on this path either.error=now shows the upstream message instead of the first 240 characters of the body. The chat routes keep their existing detail.src/copilot/responses.tshas no failed-response path; its in-stream failures stay a generic 502. The WebSearch bridge's own/responsescall is unchanged: a failure there becomes a failed search result, not a relay error. Anerrorevent inside a native stream that has already started is not inspected.tests/unit/prompt-too-long.test.tspins it.claude-haiku-4.5between 136000 and 200000 tokens (not tested, to avoid a real inference).Checklist
config.default.yaml, README, andwiki/(EN and ZH). (No config change.)After merge
🤖 Generated with Claude Code