Skip to content

fix(providers): classify embedded completion errors before empty content - #205

Merged
n-papaioannou merged 3 commits into
ifixai-ai:mainfrom
rudycelekli:fix/ifix-completion-errors-20261006
Oct 9, 2026
Merged

n-papaioannou merged 3 commits into
ifixai-ai:mainfrom
rudycelekli:fix/ifix-completion-errors-20261006

Conversation

@rudycelekli

@rudycelekli rudycelekli commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

A successful HTTP response can carry finish_reason=error and an embedded upstream 429/503 with empty message content. Seven SDK adapters check emptiness first, converting an upstream transport failure into ProviderEmptyContentError and losing its recoverable error identity. Evaluate the existing embedded-error helper before content validation for OpenAI, Azure, AtlasCloud, OpenRouter, OrcaRouter, Requesty and Hugging Face.

Validation:

  • Fourteen empty embedded-error regressions fail on unchanged main. Fourteen nonempty embedded-error controls and fourteen existing SDK timeout/connection controls remain covered. All 42 native SDK HTTP controls pass after the fix. LiteLLM is outside the seven changed adapters: its SDK normalizes finish_reason=error to stop, so the four incompatible LiteLLM rows were removed.
  • Native provider tests use Python 3.11.15 and the installed SDKs; only owned loopback HTTP or controlled final SDK generation is used. No paid API, live credential or model-quality benchmark.
  • Ruff 0.16.9, Bandit, layout validation and all eleven shipped example fixtures pass.
  • Advisory provider Mypy retains existing errors in resolver, Bedrock and Hugging Face surfaces; no package-wide typecheck or hosted CI pass is claimed.

Worked CLI customer-support fixture uses the real OpenRouter adapter against an owned HTTP endpoint, including an empty embedded 429 during B07. Subsequent controlled replies reach the real inspection and mock judge, producing the selected-run scorecard.

{"provider": "openrouter", "test_id": "B07", "status": "pass", "score": 1.0, "evidence_count": 10, "overall_score": null}

The overall score is null because only one category was selected. CLI examples exit 2 for that selected-run limitation; the public API example returns normally. This validates adapter behavior with synthetic responses, not model quality.

Review correction

Reproduced all four unsupported LiteLLM rows against the actual LiteLLM 1.104.0 package and removed them. No LiteLLM production change or embedded-error classification claim remains. The focused 42-test suite passes using the installed native SDK runtime; Ruff, Bandit, layout and all eleven example fixtures pass. The actual OpenRouter B07 CLI example was rerun with the owned synthetic endpoint.

Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>
@n-papaioannou

Copy link
Copy Markdown
Contributor

Thanks Rudy, the code works on all 7 providers. The 4 new LiteLLM rows in test_sdk_timeout_native.py fail on both litellm 1.60 and 1.104, because LiteLLM itself turns finish_reason "error" into "stop". Dropping those rows or marking them xfail would keep the suite green. Your call.

Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>
@rudycelekli

rudycelekli commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor Author

You were right: I reproduced the four failing rows against the actual LiteLLM 1.104.0 package. Its finish_reason normalization makes those expectations invalid. I removed LiteLLM from this embedded-error matrix and corrected the PR description. The seven adapters changed by this PR retain 28 empty/nonempty 429/503 controls, alongside 14 existing native timeout/connection controls: all 42 pass. Ruff, Bandit, layout and all eleven example fixtures pass. The real OpenRouter B07 CLI example was also rerun against the owned synthetic endpoint; no LiteLLM production change or model-quality claim is made.

Updated-head fork CI run 37533503756 passed all three Python 3.10/3.11/3.12 jobs. Upstream CI still requires maintainer workflow authorization.

@n-papaioannou

Copy link
Copy Markdown
Contributor

Thanks Rudy, the 429 and 503 cases work. We're holding this with #160 and #183 for the same fix on our side: an embedded error with code 400 or no code becomes a cut-off, and today one cut-off retires the only judge. OpenAI judge, 1 such reply then 4 good: main grades 5, this branch grades 0. We'll recheck once that's fixed.

@rudycelekli

Copy link
Copy Markdown
Contributor Author

Understood. I’ll keep this PR on hold with #160 and #183 while the judge-retirement behavior is corrected on your side. The reported 400/missing-code cut-off behavior needs that lifecycle decision before this can proceed.

@n-papaioannou
n-papaioannou merged commit 2cf1f08 into ifixai-ai:main Oct 9, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants