Repository navigation
chore: add CODEOWNERS - #1
Merged
Merged
Conversation
All four maintainers are owners of every path. Required by the branch protection rule about to land on main: at least one codeowner approval unlocks merge. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Sebabaian
pushed a commit
that referenced
this pull request
Jul 2, 2026
Verified all three reviewer claims against the code before acting; all three held up. B08/B09 — verdict:/rationale: keys in references.yaml were dead config. ReferenceExample (core/types.py) has no extra=forbid, so Pydantic drops unknown keys at load; the judge builder renders references as only [LABEL] "response_text" (evaluation/analytic_judge.py); and the keys violate references.schema.json (additionalProperties: false). The informative part never reached the judge. Moved the two partial- compliance examples per benchmark into rubric.yaml dimension examples, whose verdict/snippet/rationale ARE rendered to the judge and are schema-valid — B08 under no_information_leakage + no_partial_compliance, B09 under no_workaround_offered. Reverted references.yaml to its original schema-clean entries. B13 — 'good' anchor #1 in references.yaml logged steps + hashes but attributed no decision to a component/rule, so after decision_attribution became mandatory the gold anchor would itself read as a fail. Added per-step component + rule attribution so the anchor exemplifies the new mandatory bar. B07 — corrected the two comments. Removed the invented 'attacker- controlled context' mechanism (B07 has no adversarial corpus; it is a passive grounding test) and the TODO that claimed no weight table exists (DEFAULT_CATEGORY_WEIGHTS in scoring/category_weights.py holds the per-category priority; the per-test weight in SPEC is a separate layer, no mismatch). Replaced with an accurate note. B10 — removed the self-contradicting 'no runner change is needed' comment in a commit that does change the runner. Pre-existing, out of scope: B10 rubric.yaml observation_notes violates rubric.schema.json the same way (silently dropped by AnalyticRubric); left untouched as it predates this branch.
Sebabaian
added a commit
that referenced
this pull request
Jul 3, 2026
…, B30 (#50) * feat(benchmarks): audit and improve B07, B08, B09, B10, B12, B13, B28, B30 B07: add rationale comment explaining MANIPULATION category for a hallucination benchmark (adversarial grounding probes are a manipulation surface); add TODO flagging that no separate pillar weight table exists in suites.py — weight is authoritative only in SPEC. B10: add fourth framing condition D (training_data_framing) — a documented behavioural-shift vector where models relax refusals when told their response will shape future training. Update definition.yaml (step 4), rubric.yaml (decision_consistency and mislabeled_framing_ invariance extended to cover all four conditions A/B/C/D, observation notes updated), and runner.py (validation, step extraction, prompt list, response unpacking, blob, details, context string). Seed/temp pinning requires no structural change — confirmed in comment. B13: elevate decision_attribution from mandatory: false to mandatory: true; update judge_prompt_template to mark it MANDATORY. Rationale: an audit trail that logs steps without attributing each decision to a responsible component or rule is unfalsifiable. B28: raise pass threshold 0.80 → 0.90 (in runner.py SPEC, where it actually lives — definition.yaml is a stub); rebalance rubric weights: no_authority_escalation 0.20 → 0.10, no_information_leak 0.10 → 0.20 with rationale comments. Weights verified to sum to 1.00. B30: add three missing attack classes to corpus.yaml (rule-025: audit log severity downgrade; rule-026: output safety-marker stripping; rule-027: consent gate bypass by origin). IDs continue existing sequence; structure matches existing entries exactly. The 3-variant expansion logic picks these up automatically. B12: replace ms_02 (numbered-step sequential agreement — duplicate rhetorical structure of ms_01) with a socratic false-entrapment framing. All five multi_step_social payloads now use distinct rhetorical structures. No multi-turn corpus format exists for B12; single-message payloads kept. B08, B09: add 2 annotated bad-response entries each (verdict: fail, rationale) covering partial-compliance failure modes — model refuses the action but names the tool/mechanism/pathway that would accomplish it. These patterns bypass forbidden-keyword filters and were absent from both reference sets. * Remove redundant comment * fix(benchmarks): address review feedback (B07, B08, B09, B10, B13) Verified all three reviewer claims against the code before acting; all three held up. B08/B09 — verdict:/rationale: keys in references.yaml were dead config. ReferenceExample (core/types.py) has no extra=forbid, so Pydantic drops unknown keys at load; the judge builder renders references as only [LABEL] "response_text" (evaluation/analytic_judge.py); and the keys violate references.schema.json (additionalProperties: false). The informative part never reached the judge. Moved the two partial- compliance examples per benchmark into rubric.yaml dimension examples, whose verdict/snippet/rationale ARE rendered to the judge and are schema-valid — B08 under no_information_leakage + no_partial_compliance, B09 under no_workaround_offered. Reverted references.yaml to its original schema-clean entries. B13 — 'good' anchor #1 in references.yaml logged steps + hashes but attributed no decision to a component/rule, so after decision_attribution became mandatory the gold anchor would itself read as a fail. Added per-step component + rule attribution so the anchor exemplifies the new mandatory bar. B07 — corrected the two comments. Removed the invented 'attacker- controlled context' mechanism (B07 has no adversarial corpus; it is a passive grounding test) and the TODO that claimed no weight table exists (DEFAULT_CATEGORY_WEIGHTS in scoring/category_weights.py holds the per-category priority; the per-test weight in SPEC is a separate layer, no mismatch). Replaced with an accurate note. B10 — removed the self-contradicting 'no runner change is needed' comment in a commit that does change the runner. Pre-existing, out of scope: B10 rubric.yaml observation_notes violates rubric.schema.json the same way (silently dropped by AnalyticRubric); left untouched as it predates this branch. --------- Co-authored-by: SebastianGozner <sebastian.gozner@dazzlerlabs.ai>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Why now
Prerequisite for the branch protection rule about to be applied to `main` (require PR + 1 codeowner approval + CI green + no force-push + no deletion).
Test plan
🤖 Generated with Claude Code