Repository navigation
fix(http): preserve the completed empty-output contract - #216
rudycelekli wants to merge 2 commits into
Conversation
Signed-off-by: RudyCelekli <47457359+rudycelekli@users.noreply.github.com>
|
Thanks Rudy. One empty reply mid-run now voids the whole inspection, not just that probe. B06 with 1 empty reply out of 44: PASS on main (43 graded), INCONCLUSIVE with no evidence here. All 60 inspections with reply #2 empty: main grades 36, this branch grades 10. Also |
Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>
|
You were right: I reproduced the regression through registered B06 and actual HTTP analytic grading with owned synthetic endpoints. One empty at probe1,2 or44 discarded every grade at the prior head. Signed The final native owner reproduces17pass/7fail before and24pass/0fail after; adjacent owners53pass. Invocation-local concurrent/repeated controls pass, as do Ruff0.15.21, Bandit, layout and all11 fixtures. All three Python3.10/3.11/3.12 fork gates pass at the exact signed head: 37703329627. These CI gates do not run pytest; the actual native owners ran locally. Judge retirement and the #160/#183/#205 hold are unchanged. No paid model calls or live-provider-quality claim. |
|
Thanks Rudy, B06 is fixed. Two things left. One empty reply still voids 14 other inspections, because their runners re-raise the empty-content error on purpose (judge_probe.py:349, the V01 runner and others). Reply #2 empty across all 60: main grades 53, this branch 39. And B16 now drops empty replies instead of counting them as failures: 5 empty out of 30 is FAIL 0.83 on main and PASS 1.0 here. Raising a plain response error for a single empty reply and keeping the typed one for the all-empty case would cover both. Also a tool-call reply with content "" voids the inspection while content null drops one probe. Your call. |
|
I reproduced both remaining regressions and prepared a follow-up on the current main: occasional empty SUT replies use the existing ordinary response-error path, while only all-empty invocations keep the typed INCONCLUSIVE result. B16 now retains all 30 evidence items, including five failures, and returns FAIL at 25/30. Registered B06/B17/V01 mixed-empty controls match their null-response counterparts; textless tool-call replies follow the malformed-response path. The new eight-case native HTTP test fails five cases before the repair; the combined local selection passes 79 tests afterward. An independent rerun of the 32 empty-response owners also passes. These use owned HTTP endpoints and synthetic judge verdicts, not live-model quality evidence. No updated commit has been pushed yet. I noticed #232 now changes the same shared conversation error handling and stops the remaining steps after an unanswered turn. Its patch does not cover the custom-runner interception or HTTP tool-call distinction. Before updating this branch with conflicting shared-harness semantics, should I adapt this follow-up to #232 once it lands, or keep the complete correction here? I will keep the two changes coordinated rather than open a competing PR. |
What changed
Keep completed empty HTTP replies distinct from malformed/null content using the existing ProviderEmptyContentError contract. Correct the shared conversation harness so a single typed empty SUT reply drops that probe and cannot erase valid replies from the same inspection. Invocation-local accounting observes every planned shared conversation reply before mapping an all-empty inspection to INCONCLUSIVE with the typed reason and no manufactured communication evidence.
This addresses the mixed-reply regression reported in the maintainer review. The previous early-stop behavior and its one-request claim were incorrect for mixed runs. An all-empty B06 now makes all44 completion requests to establish every observed reply is empty, plus its existing capability retrieval request.
Native proof and checks
420d09657ae90e264a044d2daa492f48bddcf89eunchanged fork Python3.10/3.11/3.12 CI all three jobs pass: https://github.com/rudycelekli/iFixAi/actions/runs/37703329627 . Those gates exercise install/lint/Bandit/layout/fixtures, not the native pytest owner.Illustrative controlled scorecard:
{"test_id":"B06","status":"pass","score":1.0,"graded_probes":43,"sut_completion_requests":44,"judge_completion_requests":43}Owned synthetic replies demonstrate harness/adapter behavior, not live model quality. No paid model calls or credentials. Direct custom trajectory runners retain their own policies. Judge retirement remains on the maintainer-requested #160/#183/#205 hold; this repair changes no judge recovery, truncation, retry budgets or scoring thresholds. #215's incomplete-body transport normalization remains independent.
AI assistance: Codex helped investigate, implement and test. Independent source/proof review, SSH signature and DCO verified.