docs(benchmarks): known-answer import scoping — c-CRAB/CR-Bench TS/JS premise falsified - #307
Conversation
|
Warning Review limit reached
Next review available in: 17 minutes Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available. How can I continue?After more reviews become available, a review can be triggered using the To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews. How do review limits work?CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability. For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window. Please refer docs for additional details. Review details⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Plus Run ID: 📒 Files selected for processing (1)
📝 WalkthroughWalkthroughThe benchmark documentation removes planned c-CRAB and CR-Bench TS/JS imports. It records GHSA/npm advisory mining and SWE-Bench Multimodal transformation as proposed alternatives. Absolute-recall measurement remains pending approval and implementation. ChangesBenchmark corpus evaluation
Estimated code review effort: 1 (Trivial) | ~5 minutes Possibly related PRs
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Actionable comments posted: 2
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@docs/benchmarks/corpus-growth.md`:
- Around line 95-96: Update item 5 in the Pending work list to remove the
rejected c-CRAB/CR-Bench import request and instead require ratification and
construction of one proposed alternative, keeping the runbook consistent with
the section’s viability decision.
- Around line 88-89: Update the CR-Bench statement in the artifact comparison
paragraph to say it “has no publicly released artifact” instead of “has
published no artifact at all,” preserving the intended scoped claim without
changing the surrounding text.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: defaults
Review profile: CHILL
Plan: Pro Plus
Run ID: e98c059d-14ed-422c-9453-d2af8039704a
📒 Files selected for processing (1)
docs/benchmarks/corpus-growth.md
… fix pending-list conflicts Review round on #307: 'published no artifact at all' overclaimed (verified absence of a PUBLIC artifact only) — now 'no publicly released artifact'. Pending-work item 5 still requested the rejected c-CRAB/CR-Bench import, contradicting measurement-rule 2's viability finding — now points at ratifying one of the proposed replacements. Also marked pending item 2 (κ + noise floor) DONE per #304, same staleness class. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… premise falsified Scoped the queued absolute-recall import against the actual artifacts: c-CRAB builds on SWE-CARE and CR-Bench transforms SWE-Bench — both Python-only, so the planned "filtered to TS/JS" slice does not exist in either. c-CRAB's dataset repo additionally has no license; CR-Bench has released no artifact. Runbook item 2 now records the finding and proposes (unratified) replacements: GHSA/npm-advisory mining with fix commits for the security suites, and applying CR-Bench's transformation recipe to SWE-Bench Multimodal's JS/TS repos for correctness. Absolute recall stays blocked until one is ratified. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… fix pending-list conflicts Review round on #307: 'published no artifact at all' overclaimed (verified absence of a PUBLIC artifact only) — now 'no publicly released artifact'. Pending-work item 5 still requested the rejected c-CRAB/CR-Bench import, contradicting measurement-rule 2's viability finding — now points at ratifying one of the proposed replacements. Also marked pending item 2 (κ + noise floor) DONE per #304, same staleness class. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
c58d59a to
36d202d
Compare
…clone-gate ruling (#310) benchmarks-grow-from-telemetry gains its convergence record before the release that ships it: capture loop closed end-to-end (#295/#302/#303/#309, first 8 pure-telemetry rows, corpus 128), label-trust precondition met (#304: κ 0.735 post-triage, 4.2% noise floor; cleanlab floor still pending bench pred_probs), and the Target's c-CRAB/CR-Bench known-answer path recorded as falsified (#307) with the replacement candidates awaiting ratification. New axis clone-gate-non-import-code ([VALIDATED]): clones are measured over non-import code, excluded at the jscpd tokenizer — with the six-hole failure of post-hoc fragment classification recorded as the rejected road so a future simplifier can't silently re-vacuous the gate (#305/#308). Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Finding
The queued absolute-recall plan — "c-CRAB (arXiv 2603.23448) + CR-Bench (arXiv 2603.11078), filtered to TS/JS" — was scoped against the actual papers and released artifacts and is not buildable as specified:
github.com/c-CRAB-Benchmark/dataset) is the curation pipeline + results, with no license.Runbook change
Item 2 of the measurement-rules section now records the scoping (dated 2026-08-02) and proposes two unratified replacements for the same goal:
Absolute-recall numbers stay blocked until one is ratified and built. This is the second handover premise falsified by verification this cycle (after the rebutted-drop distribution in #306) — both now recorded where the next session will read them.
🤖 Generated with Claude Code
Summary by CodeRabbit