Repository navigation
fix: stop silently dropping --system-prompt (B26/B30/B31/B32) and fixture metadata overrides - #68
Merged
Conversation
…obes B26, B30, B31 and B32 build their probe message list inline as `[ChatMessage(role="user", ...)]` and call `provider.send_message` directly, bypassing `harness.base.send_single_turn`. That helper is what prepends `build_system_message(config.system_prompt, config.run_nonce)`, so for these four inspections the system prompt supplied via `--system-prompt` never reached the model: they scored a bare model with no deployed configuration, and the run_nonce anti-cache marker was absent. Every other inspection routes through `send_single_turn` or `run_conversation`, so a single run could grade some inspections with the system prompt and others without it. The impact is largest on B30 (malicious deployer rules) and B32 (off-topic detection): both grade refusal behaviour that a system prompt is precisely what establishes. On a real agent, B30 measured 0.68 without the prompt and 1.00 with it — the same model and the same seeds. Route all four through `send_single_turn`. B26's module-level `_INSPECTION_MESSAGE` list becomes `_INSPECTION_PROMPT`, since the message list is now built per call.
`FixtureMetadata` declares three author-facing override fields — `on_topic_examples` (B32), `b06_probes` (B06) and `case_id_prefixes` (B31) — and the inspections read them. But `_parse_fixture` builds `FixtureMetadata` from an explicit field list that omits all three, so they always fell back to their `default_factory=list` empty defaults. Setting them in a fixture had no effect. `case_id_prefixes` was already documented in the fixture JSON schema, so authors could set it, pass validation, and still get the built-in ESC/INC/TKT set. The most visible symptom is B32 on a single-tool system: it cannot derive enough on-topic prompts from one tool and errors out with a message telling the author to set `fixture.metadata.on_topic_examples` — which the loader then discards. Copy all three off the raw metadata, and document `on_topic_examples` and `b06_probes` in the fixture schema alongside `case_id_prefixes`.
stdevMac
requested review from
Sebabaian,
dimneo,
n-papaioannou and
stefyi-4355
as code owners
July 24, 2026 20:26
Contributor
Author
|
Heads-up on CI, unrelated to this PR: the lint step is already failing on I measured this branch against that same baseline: 414 before, 414 after. The apparent deltas are the same findings at shifted line numbers, because this PR is net −2 lines. So whenever the workflow gets approved, any lint failure here is inherited from Pinning ruff (e.g. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Two independent bugs where configuration the operator supplied is silently dropped before it reaches the thing that consumes it. Both are quiet: no warning, no error, and the run still produces a confident-looking score.
Found while running ifixai against a production restaurant-operations agent (Gemini 2.5 Flash + a real deployed system prompt, judge Gemini 2.5 Pro).
1.
--system-promptnever reaches B26 / B30 / B31 / B32harness.base.send_single_turnis what prependsbuild_system_message(config.system_prompt, config.run_nonce). Four inspections skip it and build the message list inline instead:So for B26, B30, B31 and B32 the model receives only the user turn. The deployed system prompt is absent, and so is the
run_nonceanti-cache marker. Every other inspection routes throughsend_single_turnorrun_conversation, so one run can grade some inspections with the system prompt and others without it — silently.This matters most for B30 (malicious deployer rules) and B32 (off-topic detection), because both grade refusal behaviour, and a system prompt is precisely what establishes refusal behaviour. An agent with an explicit "never author access-control or redaction rules" clause was still scored as if it had no instructions at all.
Measured on a real agent, same model, same seeds, only this fix changing:
The 0.68 was the harness grading a bare
gemini-2.5-flash.Fix: route all four through
send_single_turn. B26's module-level_INSPECTION_MESSAGElist becomes_INSPECTION_PROMPT, since the message list is now built per call.A note on B26: it is a rapid-fire soak probe, so including the system prompt does raise its token cost. I included it for consistency — the probe is meant to exercise the deployed configuration — but happy to drop that hunk if you would rather the load probe stay prompt-free.
2. Fixture metadata overrides are parsed and thrown away
FixtureMetadatadeclares three author-facing override fields, and the inspections read them:on_topic_examplesfixture.metadata.on_topic_examples)b06_probesgetattr(fixture.metadata, "b06_probes", None))case_id_prefixesfixture.metadata.case_id_prefixes)_parse_fixturebuildsFixtureMetadatafrom an explicit field list that omits all three, so they always fall back todefault_factory=list. Setting them in a fixture has no effect.case_id_prefixesis already documented infixtures/schema.json, so an author can set it, passifixai validate, and still silently get the built-in ESC/INC/TKT set.The sharpest symptom is B32 on a single-tool system: it cannot derive enough on-topic prompts from one tool, and errors out with a message telling the author to set
fixture.metadata.on_topic_examples— which the loader then discards. There is no way out of that loop from a fixture.Fix: copy all three off the raw metadata, and document
on_topic_examplesandb06_probesin the fixture schema next tocase_id_prefixes.Reproducing
Drop this in the repo root and run it on
main, then on this branch. It drives each affected inspection's real send path with a recording provider — no API keys, no network.verify_system_prompt.pyAnd for the loader,
_parse_fixturewith all three fields set in metadata:Checks
ifixai validate→Layout valid: 45 tests, exit 0ifixai/fixtures/examples/*.yamlvalidatebandit -r ifixai -ll→ exit 0, no new findingsruff check ifixai→ identical finding count before and after (414 → 414; the deltas are the same findings at shifted line numbers). Net −2 lines, and three now-unusedChatMessageimports removed.Notes
The two commits are independent and can be taken separately. Scores produced by B26/B30/B31/B32 before this fix are not comparable with scores after it, since the SUT configuration genuinely changes — that may be worth a line in the release notes.
Separately, I hit two robustness problems running a reasoning model (Gemini 2.5 Pro) as the judge: thinking tokens count against
max_tokens, soJUDGE_MAX_TOKENS_FLOOR = 512gets fully consumed and the verdict JSON is truncated; and OpenRouter intermittently returns HTTP 200 with empty content andfinish_reason=errorunder concurrent judge load, which aborts a whole run. I have fixes for both but kept them out of this PR since they involve tuning choices that are yours to make — happy to open a second PR if useful.