You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit bebd35b
Browse filesBrowse the repository at this point in the historyBrowse files
sync: RAG integrity, concurrency hardening, CLI polish from dev (#7)
* sync: RAG integrity, concurrency hardening, CLI polish from dev
Single squashed sync from ifixai-ai/diagnostic-dev to keep the public
release tree current with internal development.
Highlights
- RAG context integrity: B28 inspection rewritten to test prompt-injection
resistance through retrieved context, with structural typed cases.
- Judge prompt isolation: SUT response moved out of the system prompt into
a delimited user message to mitigate self-judging the response.
- Rubric cache: lazy-init the asyncio lock so multi-loop test runs don't
collide.
- Concurrency governor: ramp waiters back up gradually after a 429 instead
of releasing all parked coroutines in a thundering herd.
- Category summary: stop rendering "✓ all passed" when zero tests were
scored; show "— no scored tests" instead.
- CLI polish: per-test folder layout, benchmark progress display, run
summary terminology, PowerShell rendering fix.
- Restore iMe Core branding modules (_branding.py, _imecore_prompt.py)
and rewire run.py to use print_startup_banner and
print_imecore_conclusion, plus the --quiet flag. Public-side intent
from PR #2 preserved.
- Docs: README repo-prep, methodology trim, drop internal spec IDs from
public surface.
Sync window: public/main (deb9ecb) → diagnostic-dev/main (293a62d), 81
non-merge commits.
* feat(scorecard): introduce inconclusive status and remove canned remediation
* fix(reporting): correct inconclusive predicate, scrub recommendation surfaces, fix footer
- _print_inconclusive_summary now predicates on TestStatus.INCONCLUSIVE
per test rather than EvaluationMethod.JUDGE per evidence item. The
prior predicate counted every judge-scored evidence item including
passes, leading to a misleading "N evidence items" message.
- Lazy-init of _rubric_cache_lock in analytic_judge moved to module
scope, removing a TOCTOU window where two coroutines could each see
None and create independent locks. asyncio.Lock() at module scope is
loop-agnostic on Python >=3.10 (the project minimum).
- Drop remaining recommendation/remediation surfaces from the report:
* GRADE_INTERPRETATIONS verdict prose blockquote
* Gap Analysis section (current/required/deficit/priority blocks)
* Per-framework "Gap Details" subsections (NOT RUN coverage map)
* JSON gaps[] and grade_interpretation fields
- Footer now distinguishes package version from methodology spec
version (was rendering spec version as if it were software version).
- Category bar palette: orange/yellow/green/blue/pink.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(reporting): replace footer with iMe Core call-to-action
Replaces the version footer with the iMe Core marketing copy used
across the public surface. Drops the now-unused VERSION import.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* docs(readme): add license, python, CI, inspections, and good-first-issue badges
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
<ahref="https://github.com/ifixai-ai/diagnostic/issues?q=is%3Aopen+label%3A%22good+first+issue%22"><imgsrc="https://img.shields.io/github/issues/ifixai-ai/diagnostic/good%20first%20issue?label=good%20first%20issues&color=7057ff"alt="good first issues" /></a>
23
+
</p>
24
+
17
25
---
18
26
19
27
iFixAi runs up to 32 inspections against any AI agent and reports where its
Copy file name to clipboardExpand all lines: docs/methodology.md
+3-3Lines changed: 3 additions & 3 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -1,8 +1,8 @@
1
-
# ifixai Methodology
1
+
# iFixAi Methodology
2
2
3
-
This page states, in one read, *how*ifixai scores an AI agent and *why* each choice is defensible. It exists so a reviewer does not have to reconstruct the rules from the code.
3
+
This page states, in one read, *how*iFixAi scores an AI Agent or Deployment and *why* each choice is defensible. It exists so a reviewer does not have to reconstruct the rules from the code.
4
4
5
-
ifixai is a diagnostic, not a certification. It runs 32 inspections against any agent and reports where the deployment's response behaviour differs from common governance expectations. It is useful for CI regression tracking, vendor comparisons under a controlled fixture, and pre-audit spot checks. It is not a substitute for domain-specific threat modelling or a formal safety argument.
5
+
iFixAi is a diagnostic, not a certification. It runs 32 inspections against any agent and reports where the deployment's response behaviour differs from common governance expectations. It is useful for CI regression tracking, vendor comparisons under a controlled fixture, and pre-audit spot checks. It is not a substitute for domain-specific threat modelling or a formal safety argument.
0 commit comments