Skip to content

fix(reporting): preserve grading provenance in human exports - #201

Merged
n-papaioannou merged 2 commits into
ifixai-ai:mainfrom
rudycelekli:fix/ifix-report-review-20261006
Oct 6, 2026
Merged

n-papaioannou merged 2 commits into
ifixai-ai:mainfrom
rudycelekli:fix/ifix-report-review-20261006

Conversation

@rudycelekli

Copy link
Copy Markdown
Contributor

Human exports drop caveats which the run already records: Markdown omits warnings, the interactive scorecard omits both warning collections, and evidence rows label every passed=False as a failed answer. An instrument error or operator diagnostic can consequently appear to be a measured failure without the self-judge or invalid-run warning.

Preserve the existing warning collections in both human exports. Carry extraction_error and is_diagnostic into interactive evidence, show “grading unavailable” with its cause and “diagnostic” for operator evidence, and keep actual failed answers labelled as failures. These are provenance labels; individual inspections' scoring and mandatory gate policies are unchanged.

Validation:

  • Unchanged main fails three warning-preservation and five evidence-provenance regressions. The combined reporting, UTC and JSON audit controls pass: 26 tests.
  • Execute the production artifact JavaScript with Node against a controlled scorecard: self-judge and invalid-run warnings are visible; a budget error renders “grading unavailable” with its cause; diagnostic evidence renders “diagnostic”; the failed-answer control still renders “fail”. JSON/script-boundary escaping remains covered.
  • Required Ruff 0.16.9, Bandit, layout validation and all eleven shipped example fixtures pass.
  • Advisory scoped Mypy retains seven identical current-main scorecard diagnostics; the artifact module has none. Hosted CI is reported separately.

Illustrative export control, built using the real reporting payload/render functions, not a model-quality measurement:

{"warnings":["self-judge bias: Standard-mode score needs independent verification"],"validation_warnings":["run_invalid: controlled measurement failure; ignore grade"],"evidence":{"extraction_error":"budget","is_diagnostic":false},"rendered_badge":"grading unavailable"}

Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>
@n-papaioannou
n-papaioannou merged commit a27b238 into ifixai-ai:main Oct 6, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants