Skip to content

bench(reviewer-eval): correctness batch 2 — corpus 90 → 120 (+15 golds, +15 pairs) - #298

Merged
norvalbv merged 1 commit into
mainfrom
bench/corpus-batch-2
Aug 1, 2026
Merged

bench(reviewer-eval): correctness batch 2 — corpus 90 → 120 (+15 golds, +15 pairs)#298
norvalbv merged 1 commit into
mainfrom
bench/corpus-batch-2

Conversation

@norvalbv

@norvalbv norvalbv commented Aug 1, 2026

Copy link
Copy Markdown
Owner

PR C — stacked on #297. Second mined tranche through the propose/finalize pipeline: 15 outcome=fixed CodeRabbit findings as anonymized correctness golds + a minimal-pair decoy each. Decoys now 55 of the ~120 the correctness-reviewer-precision record names as its revisit threshold.

Deliberately appended after checkpoint evt-…-b0dc29a7d531 so the published evidence hashes the exact tree it measured; #295's row-set corpusHash keeps the 90-row baseline paired over retained rows through this append (that property, not coincidence).

Honest gap: no rebutted-thread decoys this tranche — they mostly fail the pipeline's hard drops (missing line anchors / truncated hunks). Next mining cycle should relax extraction for those threads specifically.

bench.mts validate green at 120 rows.

🤖 Generated with Claude Code

@coderabbitai

coderabbitai Bot commented Aug 1, 2026

Copy link
Copy Markdown
Contributor

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 2be4a653-d51b-47b9-b2c5-0acb39c3341a

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

…ds, +15 pairs)

Second mined tranche via the propose/finalize pipeline: 15 outcome=fixed
CodeRabbit findings adapted into anonymized correctness golds (one lens each,
caseId/sourcePr/outcomeEvidence/scopeConfirmed stamped) + a minimal-pair decoy
per gold (real fix applied, variantOf/adapted). Decoys now 55 of the ~120 the
correctness-reviewer-precision record names as its revisit threshold. Appended
AFTER checkpoint evt-...-b0dc29a7d531 (frozen 90-row tree); the row-set
corpusHash keeps the 90-row baseline paired over retained rows. No rebutted-
thread decoys this tranche: they mostly fail the hard drops (missing line
anchors / truncated hunks) — next cycle picks them up from the raw thread data.
bench validate green (120 rows); tracker views re-rendered (freshness stamps).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@norvalbv
norvalbv force-pushed the bench/corpus-batch-2 branch from 6047313 to 668c3fc Compare August 1, 2026 21:09
@norvalbv
norvalbv changed the base branch from bench/publish-correctness-90 to main August 1, 2026 21:10
@norvalbv
norvalbv merged commit 9d27442 into main Aug 1, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant