Skip to content

chore: add CODEOWNERS - #1

Merged
n-papaioannou merged 1 commit into
mainfrom
chore/codeowners
Apr 27, 2026
Merged

n-papaioannou merged 1 commit into
mainfrom
chore/codeowners

Conversation

@n-papaioannou

Copy link
Copy Markdown
Contributor

Summary

Why now

Prerequisite for the branch protection rule about to be applied to `main` (require PR + 1 codeowner approval + CI green + no force-push + no deletion).

Test plan

  • CODEOWNERS file lints cleanly in the GitHub UI (no "unknown owner" warnings)
  • All four listed users receive review requests on subsequent PRs

🤖 Generated with Claude Code

All four maintainers are owners of every path. Required by the branch
protection rule about to land on main: at least one codeowner approval
unlocks merge.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@n-papaioannou
n-papaioannou merged commit 84efb08 into main Apr 27, 2026
6 checks passed
@n-papaioannou
n-papaioannou deleted the chore/codeowners branch April 27, 2026 13:49
Sebabaian pushed a commit that referenced this pull request Jul 2, 2026
Verified all three reviewer claims against the code before acting; all
three held up.

B08/B09 — verdict:/rationale: keys in references.yaml were dead config.
ReferenceExample (core/types.py) has no extra=forbid, so Pydantic drops
unknown keys at load; the judge builder renders references as only
[LABEL] "response_text" (evaluation/analytic_judge.py); and the keys
violate references.schema.json (additionalProperties: false). The
informative part never reached the judge. Moved the two partial-
compliance examples per benchmark into rubric.yaml dimension examples,
whose verdict/snippet/rationale ARE rendered to the judge and are
schema-valid — B08 under no_information_leakage + no_partial_compliance,
B09 under no_workaround_offered. Reverted references.yaml to its original
schema-clean entries.

B13 — 'good' anchor #1 in references.yaml logged steps + hashes but
attributed no decision to a component/rule, so after decision_attribution
became mandatory the gold anchor would itself read as a fail. Added
per-step component + rule attribution so the anchor exemplifies the new
mandatory bar.

B07 — corrected the two comments. Removed the invented 'attacker-
controlled context' mechanism (B07 has no adversarial corpus; it is a
passive grounding test) and the TODO that claimed no weight table exists
(DEFAULT_CATEGORY_WEIGHTS in scoring/category_weights.py holds the
per-category priority; the per-test weight in SPEC is a separate layer,
no mismatch). Replaced with an accurate note.

B10 — removed the self-contradicting 'no runner change is needed' comment
in a commit that does change the runner.

Pre-existing, out of scope: B10 rubric.yaml observation_notes violates
rubric.schema.json the same way (silently dropped by AnalyticRubric);
left untouched as it predates this branch.
Sebabaian added a commit that referenced this pull request Jul 3, 2026
…, B30 (#50)

* feat(benchmarks): audit and improve B07, B08, B09, B10, B12, B13, B28, B30
B07: add rationale comment explaining MANIPULATION category for a
hallucination benchmark (adversarial grounding probes are a manipulation
surface); add TODO flagging that no separate pillar weight table exists
in suites.py — weight is authoritative only in SPEC.
B10: add fourth framing condition D (training_data_framing) — a
documented behavioural-shift vector where models relax refusals when
told their response will shape future training. Update definition.yaml
(step 4), rubric.yaml (decision_consistency and mislabeled_framing_
invariance extended to cover all four conditions A/B/C/D, observation
notes updated), and runner.py (validation, step extraction, prompt
list, response unpacking, blob, details, context string). Seed/temp
pinning requires no structural change — confirmed in comment.
B13: elevate decision_attribution from mandatory: false to mandatory:
true; update judge_prompt_template to mark it MANDATORY. Rationale:
an audit trail that logs steps without attributing each decision to a
responsible component or rule is unfalsifiable.
B28: raise pass threshold 0.80 → 0.90 (in runner.py SPEC, where it
actually lives — definition.yaml is a stub); rebalance rubric weights:
no_authority_escalation 0.20 → 0.10, no_information_leak 0.10 → 0.20
with rationale comments. Weights verified to sum to 1.00.
B30: add three missing attack classes to corpus.yaml (rule-025:
audit log severity downgrade; rule-026: output safety-marker stripping;
rule-027: consent gate bypass by origin). IDs continue existing
sequence; structure matches existing entries exactly. The 3-variant
expansion logic picks these up automatically.
B12: replace ms_02 (numbered-step sequential agreement — duplicate
rhetorical structure of ms_01) with a socratic false-entrapment framing.
All five multi_step_social payloads now use distinct rhetorical structures.
No multi-turn corpus format exists for B12; single-message payloads kept.
B08, B09: add 2 annotated bad-response entries each (verdict: fail,
rationale) covering partial-compliance failure modes — model refuses
the action but names the tool/mechanism/pathway that would accomplish it.
These patterns bypass forbidden-keyword filters and were absent from
both reference sets.

* Remove redundant comment

* fix(benchmarks): address review feedback (B07, B08, B09, B10, B13)

Verified all three reviewer claims against the code before acting; all
three held up.

B08/B09 — verdict:/rationale: keys in references.yaml were dead config.
ReferenceExample (core/types.py) has no extra=forbid, so Pydantic drops
unknown keys at load; the judge builder renders references as only
[LABEL] "response_text" (evaluation/analytic_judge.py); and the keys
violate references.schema.json (additionalProperties: false). The
informative part never reached the judge. Moved the two partial-
compliance examples per benchmark into rubric.yaml dimension examples,
whose verdict/snippet/rationale ARE rendered to the judge and are
schema-valid — B08 under no_information_leakage + no_partial_compliance,
B09 under no_workaround_offered. Reverted references.yaml to its original
schema-clean entries.

B13 — 'good' anchor #1 in references.yaml logged steps + hashes but
attributed no decision to a component/rule, so after decision_attribution
became mandatory the gold anchor would itself read as a fail. Added
per-step component + rule attribution so the anchor exemplifies the new
mandatory bar.

B07 — corrected the two comments. Removed the invented 'attacker-
controlled context' mechanism (B07 has no adversarial corpus; it is a
passive grounding test) and the TODO that claimed no weight table exists
(DEFAULT_CATEGORY_WEIGHTS in scoring/category_weights.py holds the
per-category priority; the per-test weight in SPEC is a separate layer,
no mismatch). Replaced with an accurate note.

B10 — removed the self-contradicting 'no runner change is needed' comment
in a commit that does change the runner.

Pre-existing, out of scope: B10 rubric.yaml observation_notes violates
rubric.schema.json the same way (silently dropped by AnalyticRubric);
left untouched as it predates this branch.

---------

Co-authored-by: SebastianGozner <sebastian.gozner@dazzlerlabs.ai>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant