Skip to content

Commit bebd35b

Browse files
sync: RAG integrity, concurrency hardening, CLI polish from dev (#7)
* sync: RAG integrity, concurrency hardening, CLI polish from dev Single squashed sync from ifixai-ai/diagnostic-dev to keep the public release tree current with internal development. Highlights - RAG context integrity: B28 inspection rewritten to test prompt-injection resistance through retrieved context, with structural typed cases. - Judge prompt isolation: SUT response moved out of the system prompt into a delimited user message to mitigate self-judging the response. - Rubric cache: lazy-init the asyncio lock so multi-loop test runs don't collide. - Concurrency governor: ramp waiters back up gradually after a 429 instead of releasing all parked coroutines in a thundering herd. - Category summary: stop rendering "✓ all passed" when zero tests were scored; show "— no scored tests" instead. - CLI polish: per-test folder layout, benchmark progress display, run summary terminology, PowerShell rendering fix. - Restore iMe Core branding modules (_branding.py, _imecore_prompt.py) and rewire run.py to use print_startup_banner and print_imecore_conclusion, plus the --quiet flag. Public-side intent from PR #2 preserved. - Docs: README repo-prep, methodology trim, drop internal spec IDs from public surface. Sync window: public/main (deb9ecb) → diagnostic-dev/main (293a62d), 81 non-merge commits. * feat(scorecard): introduce inconclusive status and remove canned remediation * fix(reporting): correct inconclusive predicate, scrub recommendation surfaces, fix footer - _print_inconclusive_summary now predicates on TestStatus.INCONCLUSIVE per test rather than EvaluationMethod.JUDGE per evidence item. The prior predicate counted every judge-scored evidence item including passes, leading to a misleading "N evidence items" message. - Lazy-init of _rubric_cache_lock in analytic_judge moved to module scope, removing a TOCTOU window where two coroutines could each see None and create independent locks. asyncio.Lock() at module scope is loop-agnostic on Python >=3.10 (the project minimum). - Drop remaining recommendation/remediation surfaces from the report: * GRADE_INTERPRETATIONS verdict prose blockquote * Gap Analysis section (current/required/deficit/priority blocks) * Per-framework "Gap Details" subsections (NOT RUN coverage map) * JSON gaps[] and grade_interpretation fields - Footer now distinguishes package version from methodology spec version (was rendering spec version as if it were software version). - Category bar palette: orange/yellow/green/blue/pink. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(reporting): replace footer with iMe Core call-to-action Replaces the version footer with the iMe Core marketing copy used across the public surface. Drops the now-unused VERSION import. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs(readme): add license, python, CI, inspections, and good-first-issue badges --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
1 parent deb9ecb commit bebd35b

40 files changed

Lines changed: 1162 additions & 863 deletions

File tree

‎README.md‎

Lines changed: 8 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -14,6 +14,14 @@
1414
<a href="CONTRIBUTING.md">Contributing</a>
1515
</p>
1616

17+
<p align="center">
18+
<a href="LICENSE"><img src="https://img.shields.io/badge/license-Apache%202.0-blue.svg" alt="license: Apache 2.0" /></a>
19+
<a href="pyproject.toml"><img src="https://img.shields.io/badge/python-3.10%2B-blue.svg" alt="python 3.10+" /></a>
20+
<a href="https://github.com/ifixai-ai/diagnostic/actions/workflows/ci.yml"><img src="https://github.com/ifixai-ai/diagnostic/actions/workflows/ci.yml/badge.svg" alt="CI" /></a>
21+
<img src="https://img.shields.io/badge/inspections-32-orange.svg" alt="32 inspections" />
22+
<a href="https://github.com/ifixai-ai/diagnostic/issues?q=is%3Aopen+label%3A%22good+first+issue%22"><img src="https://img.shields.io/github/issues/ifixai-ai/diagnostic/good%20first%20issue?label=good%20first%20issues&color=7057ff" alt="good first issues" /></a>
23+
</p>
24+
1725
---
1826

1927
iFixAi runs up to 32 inspections against any AI agent and reports where its

‎docs/methodology.md‎

Lines changed: 3 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -1,8 +1,8 @@
1-
# ifixai Methodology
1+
# iFixAi Methodology
22

3-
This page states, in one read, *how* ifixai scores an AI agent and *why* each choice is defensible. It exists so a reviewer does not have to reconstruct the rules from the code.
3+
This page states, in one read, *how* iFixAi scores an AI Agent or Deployment and *why* each choice is defensible. It exists so a reviewer does not have to reconstruct the rules from the code.
44

5-
ifixai is a diagnostic, not a certification. It runs 32 inspections against any agent and reports where the deployment's response behaviour differs from common governance expectations. It is useful for CI regression tracking, vendor comparisons under a controlled fixture, and pre-audit spot checks. It is not a substitute for domain-specific threat modelling or a formal safety argument.
5+
iFixAi is a diagnostic, not a certification. It runs 32 inspections against any agent and reports where the deployment's response behaviour differs from common governance expectations. It is useful for CI regression tracking, vendor comparisons under a controlled fixture, and pre-audit spot checks. It is not a substitute for domain-specific threat modelling or a formal safety argument.
66

77
## Evaluation paths
88

‎ifixai/cli/main.py‎

Lines changed: 19 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1,3 +1,5 @@
1+
import sys
2+
13
import click
24

35
from ifixai._version import VERSION
@@ -21,7 +23,24 @@ def ifixai_cli() -> None:
2123
ifixai_cli.add_command(compare)
2224

2325

26+
def _ensure_utf8_stdout() -> None:
27+
if sys.platform != "win32":
28+
return
29+
import io
30+
def _fix(stream): # noqa: E306
31+
enc = getattr(stream, "encoding", "") or ""
32+
if enc.lower().replace("-", "") == "utf8":
33+
return stream
34+
buf = getattr(stream, "buffer", None)
35+
if buf is None:
36+
return stream
37+
return io.TextIOWrapper(buf, encoding="utf-8", errors="replace")
38+
sys.stdout = _fix(sys.stdout)
39+
sys.stderr = _fix(sys.stderr)
40+
41+
2442
def main() -> None:
43+
_ensure_utf8_stdout()
2544
ifixai_cli()
2645

2746

‎ifixai/cli/orchestrator.py‎

Lines changed: 201 additions & 37 deletions
Original file line numberDiff line numberDiff line change
@@ -1,23 +1,27 @@
11
import os
22
import sys
3+
import threading
34
from collections import Counter
45

56
import click
67

78
from ifixai.api import run_inspections, run_single, run_strategic
89
from ifixai.core.concurrency import ConcurrencyGovernor
910
from ifixai.core.fixture_loader import load_fixture
11+
from ifixai.harness.registry import ALL_SPECS, SPEC_BY_ID
1012
from ifixai.judge.config import JudgeConfig, JudgeProviderSpec
1113
from ifixai.providers.resolver import (
1214
_PROVIDER_CREDENTIAL_ENV_VARS,
1315
detect_available_credentials,
1416
select_cross_provider_judge,
1517
)
18+
from ifixai.scoring.category_weights import STRATEGIC_TEST_IDS
1619
from ifixai.core.types import (
1720
TestResult,
1821
TestRunResult,
19-
EvaluationMethod,
2022
EvaluationPipelineConfig,
23+
InspectionCategory,
24+
TestStatus,
2125
)
2226

2327

@@ -157,53 +161,184 @@ def _build_judge_config(
157161

158162

159163
def _print_inconclusive_summary(result: TestRunResult) -> None:
160-
inconclusive_count = 0
161-
inconclusive_tests: set[str] = set()
162-
by_category: Counter[str] = Counter()
163-
164-
for br in result.test_results:
165-
for evidence in br.evidence:
166-
if evidence.evaluation_method == EvaluationMethod.JUDGE:
167-
inconclusive_count += 1
168-
inconclusive_tests.add(br.test_id)
169-
by_category[br.category.value] += 1
170-
171-
if inconclusive_count == 0:
172-
click.echo(click.style("Inconclusive: 0 evidence items.", fg="green"))
173-
else:
174-
breakdown = ", ".join(f"{cat}={n}" for cat, n in by_category.most_common())
175-
click.echo(
176-
click.style(
177-
f"Inconclusive: {inconclusive_count} evidence items across "
178-
f"{len(inconclusive_tests)} tests ({breakdown})",
179-
fg="yellow",
180-
)
164+
inconclusive = [
165+
br for br in result.test_results
166+
if br.status == TestStatus.INCONCLUSIVE
167+
]
168+
169+
if not inconclusive:
170+
click.echo(click.style("Inconclusive: 0 tests.", fg="green"))
171+
return
172+
173+
by_category: Counter[str] = Counter(br.category.value for br in inconclusive)
174+
breakdown = ", ".join(f"{cat}={n}" for cat, n in by_category.most_common())
175+
click.echo(
176+
click.style(
177+
f"Inconclusive: {len(inconclusive)} tests ({breakdown})",
178+
fg="yellow",
181179
)
180+
)
182181

183182

184183
def _print_insufficient_evidence_summary(result: TestRunResult) -> None:
185184
insufficient = [br for br in result.test_results if br.insufficient_evidence]
186185
total = len(result.test_results)
187186
if not insufficient:
188-
click.echo(click.style(f"0 out of {total} tests have failed.", fg="green"))
187+
click.echo(
188+
click.style(
189+
f"All {total} tests produced sufficient evidence to be scored.",
190+
fg="green",
191+
)
192+
)
189193
return
190194
inspection_ids = ", ".join(sorted(br.test_id for br in insufficient))
191195
click.echo(
192196
click.style(
193-
f"{len(insufficient)} out of {total} tests have failed "
194-
f"({inspection_ids}). Wrap your provider in a governance layer "
195-
f"or run with ≥2 provider credentials. See docs/methodology.md.",
197+
f"{len(insufficient)} out of {total} tests had insufficient evidence "
198+
f"to be scored ({inspection_ids}). The remaining tests were scored but "
199+
f"may still be below threshold -- see the per-category bars above. "
200+
f"Wrap your provider in a governance layer or run with ≥2 provider "
201+
f"credentials. See docs/methodology.md.",
196202
fg="yellow",
197203
)
198204
)
199205

200206

201-
def _progress_callback(
207+
_CATEGORY_BAR_COLOR: dict[str, str] = {
208+
InspectionCategory.FABRICATION.value: "\033[38;5;208m", # orange
209+
InspectionCategory.MANIPULATION.value: "\033[93m", # yellow
210+
InspectionCategory.DECEPTION.value: "\033[92m", # green
211+
InspectionCategory.UNPREDICTABILITY.value: "\033[94m", # blue
212+
InspectionCategory.OPACITY.value: "\033[38;5;213m", # pink
213+
}
214+
_RESET = "\033[0m"
215+
_RED = "\033[91m"
216+
_GREEN = "\033[92m"
217+
_YELLOW = "\033[93m"
218+
_DIM = "\033[2m"
219+
_BOLD = "\033[1m"
220+
_SPINNER_FRAMES = ["⠋", "⠙", "⠹", "⠸", "⠼", "⠴", "⠦", "⠧", "⠇", "⠏"]
221+
222+
223+
class BenchmarkProgressDisplay:
224+
"""Live animated display: pre-prints all benchmarks then updates in-place."""
225+
226+
def __init__(self, tests: list[tuple[str, str]]) -> None:
227+
self._tests = tests # [(test_id, name), ...]
228+
self._results: dict[str, TestResult] = {}
229+
self._frame_idx = 0
230+
self._lock = threading.Lock()
231+
self._done = threading.Event()
232+
self._thread: threading.Thread | None = None
233+
234+
def start(self) -> None:
235+
for test_id, name in self._tests:
236+
sys.stdout.write(f" {_YELLOW}⠋{_RESET} {_DIM}{test_id}{_RESET} {name}\n")
237+
sys.stdout.flush()
238+
self._thread = threading.Thread(target=self._animate, daemon=True)
239+
self._thread.start()
240+
241+
def update(self, test_id: str, index: int, total: int, result: TestResult) -> None:
242+
with self._lock:
243+
self._results[test_id] = result
244+
245+
def stop(self) -> None:
246+
self._done.set()
247+
if self._thread:
248+
self._thread.join(timeout=2.0)
249+
self._redraw(final=True)
250+
251+
def _animate(self) -> None:
252+
while not self._done.wait(timeout=0.1):
253+
self._frame_idx = (self._frame_idx + 1) % len(_SPINNER_FRAMES)
254+
self._redraw()
255+
256+
def _build_lines(self, final: bool = False) -> list[str]:
257+
lines: list[str] = []
258+
with self._lock:
259+
frame = _SPINNER_FRAMES[self._frame_idx]
260+
for test_id, name in self._tests:
261+
if test_id in self._results:
262+
result = self._results[test_id]
263+
if result.insufficient_evidence:
264+
icon = f"{_YELLOW}⊘{_RESET}"
265+
status = f"{_YELLOW}INCONCLUSIVE{_RESET}"
266+
lines.append(
267+
f" {icon} {_BOLD}{test_id}{_RESET} {name} "
268+
f"... {status} (insufficient evidence)"
269+
)
270+
continue
271+
if result.passing:
272+
icon = f"{_GREEN}✓{_RESET}"
273+
status = f"{_GREEN}PASS{_RESET}"
274+
else:
275+
icon = f"{_RED}✗{_RESET}"
276+
status = f"{_RED}FAIL{_RESET}"
277+
lines.append(
278+
f" {icon} {_BOLD}{test_id}{_RESET} {name} "
279+
f"... {status} ({result.score:.0%})"
280+
)
281+
else:
282+
spinner = "·" if final else frame
283+
lines.append(
284+
f" {_YELLOW}{spinner}{_RESET} {_DIM}{test_id}{_RESET} {name}"
285+
)
286+
return lines
287+
288+
def _redraw(self, final: bool = False) -> None:
289+
n = len(self._tests)
290+
if n == 0:
291+
return
292+
lines = self._build_lines(final=final)
293+
sys.stdout.write(f"\033[{n}A")
294+
for line in lines:
295+
sys.stdout.write(f"\r\033[K{line}\n")
296+
sys.stdout.flush()
297+
298+
299+
def _print_category_summary(result: TestRunResult) -> None:
300+
if not result.category_scores:
301+
return
302+
bar_width = 22
303+
click.echo()
304+
for cs in result.category_scores:
305+
color = _CATEGORY_BAR_COLOR.get(cs.category.value, "\033[96m")
306+
bar = color + ("█" * bar_width) + _RESET
307+
total_in_suite = len(cs.test_ids)
308+
failed = cs.test_count - cs.tests_passed
309+
inconclusive = total_in_suite - cs.test_count
310+
count_str = f"{cs.test_count}/{total_in_suite}"
311+
if total_in_suite == 0:
312+
count_str = "—"
313+
fail_str = f"{_DIM}not in this suite{_RESET}"
314+
elif cs.test_count == 0:
315+
fail_str = f"{_YELLOW}⊘ {inconclusive} inconclusive{_RESET}"
316+
elif failed > 0 and inconclusive > 0:
317+
fail_str = f"{_RED}× {failed} failed{_RESET}, {_YELLOW}⊘ {inconclusive} inconclusive{_RESET}"
318+
elif failed > 0:
319+
fail_str = f"{_RED}× {failed} failed{_RESET}"
320+
elif inconclusive > 0:
321+
fail_str = f"{_GREEN}✓ {cs.tests_passed} passed{_RESET}, {_YELLOW}⊘ {inconclusive} inconclusive{_RESET}"
322+
else:
323+
fail_str = f"{_GREEN}✓ all passed{_RESET}"
324+
name = cs.category.value.ljust(16)
325+
click.echo(f" {name} [{bar}] {count_str:>5} {fail_str}")
326+
click.echo()
327+
328+
329+
def _progress_callback_plain(
202330
bid: str,
203331
index: int,
204332
total: int,
205333
bench_result: TestResult,
206334
) -> None:
335+
"""Fallback used when stdout is not a TTY (e.g. piped/redirected)."""
336+
if bench_result.insufficient_evidence:
337+
click.echo(
338+
f" [{index}/{total}] {bid} {bench_result.name} ... "
339+
f"{click.style('INCONCLUSIVE', fg='yellow')} (insufficient evidence)"
340+
)
341+
return
207342
status_label = (
208343
click.style("PASS", fg="green")
209344
if bench_result.passing
@@ -215,6 +350,20 @@ def _progress_callback(
215350
)
216351

217352

353+
def _build_display_tests(
354+
strategic: bool,
355+
test_id: str | None,
356+
) -> list[tuple[str, str]]:
357+
if test_id:
358+
uid = test_id.upper()
359+
spec = SPEC_BY_ID.get(uid)
360+
return [(uid, spec.name if spec else uid)] # type: ignore[union-attr]
361+
if strategic:
362+
strategic_set = set(STRATEGIC_TEST_IDS)
363+
return [(s.test_id, s.name) for s in ALL_SPECS if s.test_id in strategic_set]
364+
return [(s.test_id, s.name) for s in ALL_SPECS]
365+
366+
218367
async def execute_tests(
219368
provider: str,
220369
api_key: str,
@@ -242,7 +391,15 @@ async def execute_tests(
242391
click.echo(click.style(f"Fixture error: {exc}", fg="red"))
243392
return None
244393

245-
effective_callback = progress_callback or _progress_callback
394+
use_display = progress_callback is None and sys.stdout.isatty()
395+
display: BenchmarkProgressDisplay | None = None
396+
397+
if use_display:
398+
display = BenchmarkProgressDisplay(_build_display_tests(strategic, test_id))
399+
display.start()
400+
effective_callback = display.update
401+
else:
402+
effective_callback = progress_callback or _progress_callback_plain
246403

247404
try:
248405
if test_id:
@@ -260,15 +417,18 @@ async def execute_tests(
260417
sut_temperature=sut_temperature,
261418
sut_seed=sut_seed,
262419
)
263-
status_label = (
264-
click.style("PASS", fg="green")
265-
if single_result.passing
266-
else click.style("FAIL", fg="red")
267-
)
268-
click.echo(
269-
f" [1/1] {test_id} {single_result.name} ... "
270-
f"{status_label} ({single_result.score:.0%})"
271-
)
420+
if display:
421+
display.update(test_id, 1, 1, single_result)
422+
else:
423+
status_label = (
424+
click.style("PASS", fg="green")
425+
if single_result.passing
426+
else click.style("FAIL", fg="red")
427+
)
428+
click.echo(
429+
f" [1/1] {test_id} {single_result.name} ... "
430+
f"{status_label} ({single_result.score:.0%})"
431+
)
272432
return TestRunResult(
273433
system_name=system_name,
274434
system_version=system_version,
@@ -324,3 +484,7 @@ async def execute_tests(
324484
except Exception as exc:
325485
click.echo(click.style(f"Test execution failed: {exc}", fg="red"))
326486
return None
487+
488+
finally:
489+
if display:
490+
display.stop()

‎ifixai/cli/reports.py‎

Lines changed: 4 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -28,12 +28,14 @@ def save_reports(
2828
fixture_slug = _slugify(result.fixture_name)
2929
base_name = f"ifixai-{system_slug}-{fixture_slug}"
3030

31+
click.echo(click.style("Access your Full Report here:", bold=True))
32+
3133
if report_format in ("json", "both"):
3234
json_path = out_path / f"{base_name}.json"
3335
json_path.write_text(generate_json_report(result), encoding="utf-8")
34-
click.echo(f" JSON report: {json_path}")
36+
click.echo(f" JSON report: {json_path}")
3537

3638
if report_format in ("markdown", "both"):
3739
md_path = out_path / f"{base_name}.md"
3840
md_path.write_text(generate_markdown_report(result), encoding="utf-8")
39-
click.echo(f" Markdown report: {md_path}")
41+
click.echo(f" Markdown report: {md_path}")

0 commit comments

Comments
 (0)