Skip to content

fix(http): preserve the completed empty-output contract - #216

Open
rudycelekli wants to merge 2 commits into
ifixai-ai:mainfrom
rudycelekli:fix/ifix-http-empty-content-20261006
Open

rudycelekli wants to merge 2 commits into
ifixai-ai:mainfrom
rudycelekli:fix/ifix-http-empty-content-20261006

Conversation

@rudycelekli

@rudycelekli rudycelekli commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

What changed

Keep completed empty HTTP replies distinct from malformed/null content using the existing ProviderEmptyContentError contract. Correct the shared conversation harness so a single typed empty SUT reply drops that probe and cannot erase valid replies from the same inspection. Invocation-local accounting observes every planned shared conversation reply before mapping an all-empty inspection to INCONCLUSIVE with the typed reason and no manufactured communication evidence.

This addresses the mixed-reply regression reported in the maintainer review. The previous early-stop behavior and its one-request claim were incorrect for mixed runs. An all-empty B06 now makes all44 completion requests to establish every observed reply is empty, plus its existing capability retrieval request.

Native proof and checks

  • Actual registered B06 through public api.run_single, HttpProvider and native HTTP analytic grading against owned loopback endpoints. One empty at probe1/2/44 retains43 real graded probes and PASS1.0; a synthetic failing judge retains FAIL0.0. The null middle-reply control keeps43 grades and its existing malformed-response communication evidence. All-null keeps44 communication errors; it is not classified as valid empty text.
  • All-empty retains INCONCLUSIVE, the typed empty-content reason,0 communication evidence and0 judge calls after44 completion replies. Concurrent all-empty/mixed runs and subsequent registered execution confirm invocation-local accounting/reset.
  • Same final owner unchanged228da7:17pass/7fail; corrected24pass/0fail,0skip. Combined focused adjacent provider, pipeline-budget and health owners53pass/0fail. Existing14 provider transport shape controls remain unchanged, including cutoff/embedded error/ordinary text/retry receipt assertions.
  • Whole-package Ruff0.15.21, Bandit, layout60inspections and all11 shipped example fixtures pass. Advisory Mypy is blocked by an inherited NumPy stub syntax error in the reused cross-version dependency graph; no full typecheck pass is claimed.
  • Exact signed-head 420d09657ae90e264a044d2daa492f48bddcf89e unchanged fork Python3.10/3.11/3.12 CI all three jobs pass: https://github.com/rudycelekli/iFixAi/actions/runs/37703329627 . Those gates exercise install/lint/Bandit/layout/fixtures, not the native pytest owner.

Illustrative controlled scorecard:

{"test_id":"B06","status":"pass","score":1.0,"graded_probes":43,"sut_completion_requests":44,"judge_completion_requests":43}

Owned synthetic replies demonstrate harness/adapter behavior, not live model quality. No paid model calls or credentials. Direct custom trajectory runners retain their own policies. Judge retirement remains on the maintainer-requested #160/#183/#205 hold; this repair changes no judge recovery, truncation, retry budgets or scoring thresholds. #215's incomplete-body transport normalization remains independent.

AI assistance: Codex helped investigate, implement and test. Independent source/proof review, SSH signature and DCO verified.

Signed-off-by: RudyCelekli <47457359+rudycelekli@users.noreply.github.com>
@n-papaioannou

Copy link
Copy Markdown
Contributor

Thanks Rudy. One empty reply mid-run now voids the whole inspection, not just that probe. B06 with 1 empty reply out of 44: PASS on main (43 graded), INCONCLUSIVE with no evidence here. All 60 inspections with reply #2 empty: main grades 36, this branch grades 10. Also content: "" voids the inspection but content: null drops one probe. Limiting it to "every reply empty" would keep the dead-agent win. Your call.

Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>
@rudycelekli

Copy link
Copy Markdown
Contributor Author

You were right: I reproduced the regression through registered B06 and actual HTTP analytic grading with owned synthetic endpoints. One empty at probe1,2 or44 discarded every grade at the prior head. Signed 420d09657ae90e264a044d2daa492f48bddcf89e now continues the planned shared-conversation probes and retains43 grades; both passing and failing grading controls keep their outcomes. Null content retains its existing malformed-response path. Only an all-empty shared conversation inspection receives the typed empty-content INCONCLUSIVE result. This requires observing all44 B06 completion replies, so I corrected the previous one-request claim in the body.

The final native owner reproduces17pass/7fail before and24pass/0fail after; adjacent owners53pass. Invocation-local concurrent/repeated controls pass, as do Ruff0.15.21, Bandit, layout and all11 fixtures. All three Python3.10/3.11/3.12 fork gates pass at the exact signed head: 37703329627. These CI gates do not run pytest; the actual native owners ran locally. Judge retirement and the #160/#183/#205 hold are unchanged. No paid model calls or live-provider-quality claim.

@n-papaioannou

Copy link
Copy Markdown
Contributor

Thanks Rudy, B06 is fixed. Two things left. One empty reply still voids 14 other inspections, because their runners re-raise the empty-content error on purpose (judge_probe.py:349, the V01 runner and others). Reply #2 empty across all 60: main grades 53, this branch 39. And B16 now drops empty replies instead of counting them as failures: 5 empty out of 30 is FAIL 0.83 on main and PASS 1.0 here. Raising a plain response error for a single empty reply and keeping the typed one for the all-empty case would cover both. Also a tool-call reply with content "" voids the inspection while content null drops one probe. Your call.

@rudycelekli

Copy link
Copy Markdown
Contributor Author

I reproduced both remaining regressions and prepared a follow-up on the current main: occasional empty SUT replies use the existing ordinary response-error path, while only all-empty invocations keep the typed INCONCLUSIVE result. B16 now retains all 30 evidence items, including five failures, and returns FAIL at 25/30. Registered B06/B17/V01 mixed-empty controls match their null-response counterparts; textless tool-call replies follow the malformed-response path.

The new eight-case native HTTP test fails five cases before the repair; the combined local selection passes 79 tests afterward. An independent rerun of the 32 empty-response owners also passes. These use owned HTTP endpoints and synthetic judge verdicts, not live-model quality evidence. No updated commit has been pushed yet.

I noticed #232 now changes the same shared conversation error handling and stops the remaining steps after an unanswered turn. Its patch does not cover the custom-runner interception or HTTP tool-call distinction. Before updating this branch with conflicting shared-harness semantics, should I adapt this follow-up to #232 once it lands, or keep the complete correction here? I will keep the two changes coordinated rather than open a competing PR.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants