Repository navigation
fix(providers): reject native Gemini and Bedrock judge cutoffs - #183
rudycelekli wants to merge 3 commits into
Conversation
Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>
Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>
|
Thanks Rudy, this works and matches what the HTTP and OpenRouter judges already do. We're holding it for a fix on our side first: right now one cut-off judge reply retires that judge for the rest of the run, so with no fallback every later probe goes unscored. Once that's fixed on main we'll merge this. |
|
Understood. Keeping this PR on the requested hold until main's judge recovery handles a cut-off reply without retiring the only judge for the remainder of the run. Current main |
Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>
|
Merged current main into this branch and resolved the Gemini conflict within the existing request-owned client loop. The cutoff checks remain the only provider behavior change; this does not alter judge recovery or retirement. All 19 native Python 3.11 cutoff, credential-lifecycle and transport controls pass, together with Ruff, Bandit, layout and all eleven example fixtures. The unchanged fork CI passes on Python 3.10–3.12 at signed head Keeping the requested hold until your recovery fix lands on main. |
Gemini and Bedrock currently return partial judge text even when
reject_truncated=True. This patch reads their native MAX_TOKENS/max_tokens stop signals and raises the existing ProviderTruncatedError before the empty-content check. Cut-off system-under-test text remains available when rejection is disabled.Validation:
object.converseattr-defined error; no full-package typecheck or hosted CI pass is claimed.Worked synthetic fixture example: B07 through Bedrock made 11 native SDK HTTP requests; Gemini made 11 SDK-decoded calls through a controlled transport. Both retained partial SUT replies even with the native max-token stop signal, and each produced this selected-run scorecard:
{"test_id":"B07","status":"pass","score":1.0,"evidence_count":10,"overall_score":null}The judge is mock, and the overall score is null because only one category was selected (CLI exit 2). These examples validate adapter behavior, not model quality. No paid APIs or live credentials were used.