Skip to content

fix(providers): reject native Gemini and Bedrock judge cutoffs - #183

Open
rudycelekli wants to merge 3 commits into
ifixai-ai:mainfrom
rudycelekli:fix/ifix-native-judge-cutoff-20261005
Open

rudycelekli wants to merge 3 commits into
ifixai-ai:mainfrom
rudycelekli:fix/ifix-native-judge-cutoff-20261005

Conversation

@rudycelekli

Copy link
Copy Markdown
Contributor

Gemini and Bedrock currently return partial judge text even when reject_truncated=True. This patch reads their native MAX_TOKENS/max_tokens stop signals and raises the existing ProviderTruncatedError before the empty-content check. Cut-off system-under-test text remains available when rejection is disabled.

Validation:

  • Four regressions fail on unchanged main (partial and empty cutoffs for both providers), while four completed/SUT controls pass. All eight pass afterward.
  • Bedrock requests traverse the installed SDK and an owned local HTTP server. Gemini uses the installed GenerativeModel and native protobuf response conversion with a synthetic outbound transport.
  • Whole-package Ruff, Bandit, layout validation (60 inspections), and all 11 example fixtures pass.
  • Advisory focused Mypy retains an existing Bedrock object.converse attr-defined error; no full-package typecheck or hosted CI pass is claimed.

Worked synthetic fixture example: B07 through Bedrock made 11 native SDK HTTP requests; Gemini made 11 SDK-decoded calls through a controlled transport. Both retained partial SUT replies even with the native max-token stop signal, and each produced this selected-run scorecard:

{"test_id":"B07","status":"pass","score":1.0,"evidence_count":10,"overall_score":null}

The judge is mock, and the overall score is null because only one category was selected (CLI exit 2). These examples validate adapter behavior, not model quality. No paid APIs or live credentials were used.

Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>
Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>
@n-papaioannou

Copy link
Copy Markdown
Contributor

Thanks Rudy, this works and matches what the HTTP and OpenRouter judges already do. We're holding it for a fix on our side first: right now one cut-off judge reply retires that judge for the rest of the run, so with no fallback every later probe goes unscored. Once that's fixed on main we'll merge this.

@rudycelekli

Copy link
Copy Markdown
Contributor Author

Understood. Keeping this PR on the requested hold until main's judge recovery handles a cut-off reply without retiring the only judge for the remainder of the run. Current main e08f1ca8f481dff49f15f89c6df600920d438a2f still retires a model after JudgeTransportExhaustedError; no changes here override that decision. The provider cutoff patch can be revalidated after the upstream recovery fix.

Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>
@rudycelekli

Copy link
Copy Markdown
Contributor Author

Merged current main into this branch and resolved the Gemini conflict within the existing request-owned client loop. The cutoff checks remain the only provider behavior change; this does not alter judge recovery or retirement.

All 19 native Python 3.11 cutoff, credential-lifecycle and transport controls pass, together with Ruff, Bandit, layout and all eleven example fixtures. The unchanged fork CI passes on Python 3.10–3.12 at signed head 29eb8eea7021907f6a07c2d11afd3929cc22f931: https://github.com/rudycelekli/iFixAi/actions/runs/37452723496. Upstream authorization and review remain separate.

Keeping the requested hold until your recovery fix lands on main.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants