Repository navigation
fix(providers): classify embedded completion errors before empty content - #205
n-papaioannou merged 3 commits into
Conversation
Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>
|
Thanks Rudy, the code works on all 7 providers. The 4 new LiteLLM rows in test_sdk_timeout_native.py fail on both litellm 1.60 and 1.104, because LiteLLM itself turns finish_reason "error" into "stop". Dropping those rows or marking them xfail would keep the suite green. Your call. |
Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>
|
You were right: I reproduced the four failing rows against the actual LiteLLM 1.104.0 package. Its finish_reason normalization makes those expectations invalid. I removed LiteLLM from this embedded-error matrix and corrected the PR description. The seven adapters changed by this PR retain 28 empty/nonempty 429/503 controls, alongside 14 existing native timeout/connection controls: all 42 pass. Ruff, Bandit, layout and all eleven example fixtures pass. The real OpenRouter B07 CLI example was also rerun against the owned synthetic endpoint; no LiteLLM production change or model-quality claim is made. Updated-head fork CI run 37533503756 passed all three Python 3.10/3.11/3.12 jobs. Upstream CI still requires maintainer workflow authorization. |
|
Thanks Rudy, the 429 and 503 cases work. We're holding this with #160 and #183 for the same fix on our side: an embedded error with code 400 or no code becomes a cut-off, and today one cut-off retires the only judge. OpenAI judge, 1 such reply then 4 good: main grades 5, this branch grades 0. We'll recheck once that's fixed. |
A successful HTTP response can carry finish_reason=error and an embedded upstream 429/503 with empty message content. Seven SDK adapters check emptiness first, converting an upstream transport failure into ProviderEmptyContentError and losing its recoverable error identity. Evaluate the existing embedded-error helper before content validation for OpenAI, Azure, AtlasCloud, OpenRouter, OrcaRouter, Requesty and Hugging Face.
Validation:
Worked CLI customer-support fixture uses the real OpenRouter adapter against an owned HTTP endpoint, including an empty embedded 429 during B07. Subsequent controlled replies reach the real inspection and mock judge, producing the selected-run scorecard.
{"provider": "openrouter", "test_id": "B07", "status": "pass", "score": 1.0, "evidence_count": 10, "overall_score": null}The overall score is null because only one category was selected. CLI examples exit 2 for that selected-run limitation; the public API example returns normally. This validates adapter behavior with synthetic responses, not model quality.
Review correction
Reproduced all four unsupported LiteLLM rows against the actual LiteLLM 1.104.0 package and removed them. No LiteLLM production change or embedded-error classification claim remains. The focused 42-test suite passes using the installed native SDK runtime; Ruff, Bandit, layout and all eleven example fixtures pass. The actual OpenRouter B07 CLI example was rerun with the owned synthetic endpoint.