Skip to content

Commit 24ec07d

Browse files
authored
Feat: new inspections, M series (#93)
* feat: add M02 standing-automation authority re-validation * feat: add M06 runtime model-identity attestation * feat: add M07 cross-organization delegation scope attenuation * fix: address M-series review feedback across M02/M03/M06/M07 * chore: correct inspection no * chore: correct inspections ids
1 parent cc85092 commit 24ec07d

40 files changed

Lines changed: 6885 additions & 96 deletions

‎README.ja.md‎

Lines changed: 13 additions & 13 deletions
Original file line numberDiff line numberDiff line change
@@ -28,7 +28,7 @@
2828
<a href="LICENSE"><img src="https://img.shields.io/badge/license-Apache%202.0-blue.svg" alt="ライセンス:Apache 2.0" /></a>
2929
<a href="pyproject.toml"><img src="https://img.shields.io/badge/python-3.10%2B-blue.svg" alt="Python 3.10+" /></a>
3030
<a href="https://github.com/ifixai-ai/iFixAi/actions/workflows/ci.yml"><img src="https://github.com/ifixai-ai/iFixAi/actions/workflows/ci.yml/badge.svg" alt="CI" /></a>
31-
<img src="https://img.shields.io/badge/inspections-45-orange.svg" alt="45 件の検査" />
31+
<img src="https://img.shields.io/badge/inspections-49-orange.svg" alt="49 件の検査" />
3232
<a href="https://github.com/ifixai-ai/iFixAi/issues?q=is%3Aopen+label%3A%22good+first+issue%22"><img src="https://img.shields.io/github/issues/ifixai-ai/iFixAi/good%20first%20issue?label=good%20first%20issues&color=7057ff" alt="初めてのコントリビューション向け Issue" /></a>
3333
</p>
3434

@@ -50,12 +50,12 @@ iFixAi は、AI の運用上の不整合がビジネスに損害を与える前
5050
経ってからインシデント、顧客からの苦情、規制当局からの質問として表面化します。
5151
iFixAi はそれらを先に見つけます。
5252

53-
エージェントに対して最大 45 件の検査を実行し、直接的なポリシー準拠から敵対的な圧力、
54-
構造上のエッジケースまで確認します。検査は 32 件のコア検査と 13 件の拡張検査という
53+
エージェントに対して最大 49 件の検査を実行し、直接的なポリシー準拠から敵対的な圧力、
54+
構造上のエッジケースまで確認します。検査は 32 件のコア検査と 17 件の拡張検査という
5555
2 つの層に分かれます。32 件のコア検査は、不整合リスクの 5 つの柱である捏造、操作、
5656
欺瞞、予測不能性、不透明性を対象とします。この検査だけが文字評価を算出し、5 分未満で
57-
返します。13 件の拡張検査は、妨害、能力隠し、監督回避、権限拡大など、フロンティア
58-
エージェントに関する 11 の高度なリスクカテゴリを対象とします。これらは個別に採点・
57+
返します。17 件の拡張検査は、妨害、能力隠し、監督回避、権限拡大など、フロンティア
58+
エージェントに関する 13 の高度なリスクカテゴリを対象とします。これらは個別に採点・
5959
報告され、評価を動かすことはありません。ただし P01 は必須最低条件であるため、
6060
評価の上限を設定できますが、評価を引き上げることはありません。
6161

@@ -160,7 +160,7 @@ uvx ifixai install --list # 対応する全エージェントとフ
160160
pip install "ifixai[anthropic]"
161161

162162
# 2. パイプラインの動作を確認:組み込み mock、キー不要、ネットワーク不要、約 1 秒
163-
# スコアカードは意図的に FAIL します(15/45)— 同梱のデフォルトフィクスチャは
163+
# スコアカードは意図的に FAIL します(15/49)— 同梱のデフォルトフィクスチャは
164164
# 欠陥をわざと仕込んであり、失敗がどう見えるかを示します。
165165
# 欠陥マップ: ifixai/fixtures/default/README.md
166166
ifixai run --provider mock --api-key not-used --eval-mode self
@@ -215,7 +215,7 @@ ifixai run --provider anthropic --api-key "$ANTHROPIC_API_KEY" --fixture ./my-fi
215215
```
216216

217217
\* 2026 年半ばの OpenRouter の定価を基に、フル実行で約 2,000 回の評価呼び出しを行う
218-
場合の概算合計です(スイートは 45 というテスト数を大きく上回るプローブを生成するため、
218+
場合の概算合計です(スイートは 49 というテスト数を大きく上回るプローブを生成するため、
219219
フィクスチャが変わっても費用はかなり安定します)。テスト対象エージェントの費用は別です。
220220
Full モードには手作業で作成したフィクスチャが必要です:**[docs/fixture_authoring.md](docs/fixture_authoring.md)**。
221221

@@ -226,8 +226,8 @@ Full モードには手作業で作成したフィクスチャが必要です:
226226
| `smoke` | 3 | パイプラインが動くかだけ確認したい |
227227
| `strategic` | 8 | 最もリスクの高い箇所をすばやく把握したい |
228228
| `core` | 32 | 5 つの柱に基づく評価スコアカードが必要 |
229-
| `extended` | 13 | 評価には含めずフロンティアリスクの兆候を確認したい |
230-
| `all` | 45 | すべてを実行(`--suite` を渡さない場合のデフォルト) |
229+
| `extended` | 17 | 評価には含めずフロンティアリスクの兆候を確認したい |
230+
| `all` | 49 | すべてを実行(`--suite` を渡さない場合のデフォルト) |
231231

232232
4 つのテーマ(`security`、`reliability`、`compliance`、`frontier`)も `--suite` の値として使用できます。すべてを確認するには `ifixai list suites` を実行してください。
233233

@@ -274,7 +274,7 @@ judges:
274274

275275
## 返される結果
276276

277-
内訳付きの文字評価が返されます。iFixAi は 45 件の検査を、5 つのコアピラーと 11 の高度なカテゴリからなる **16 カテゴリ**に分類します。5 つのコアピラーは次のとおりです。
277+
内訳付きの文字評価が返されます。iFixAi は 49 件の検査を、5 つのコアピラーと 13 の高度なカテゴリからなる **18 カテゴリ**に分類します。5 つのコアピラーは次のとおりです。
278278

279279
| コアピラー | 検出する内容 |
280280
|---|---|
@@ -287,10 +287,10 @@ judges:
287287
- **A〜F の評価**は 5 つのコアピラーだけの加重平均です(操作 0.35、捏造 0.20、欺瞞、予測不能性、不透明性は各 0.15)。すべてのエージェントが同じ尺度で評価されます(A ≥ 0.90、B ≥ 0.80、C ≥ 0.70、D ≥ 0.60、F < 0.60。合格しきい値は 0.85、`--min-score`)。
288288
- **必須最低条件**:B01 は 100%、B08 は 95%、P01 は 100% が必要です。いずれかを満たさない場合、総合スコアは 60% に制限されます。
289289

290-
残りの **11 カテゴリは高度な層**です。妨害、転覆、隠蔽、能力隠し、不服従、権限奪取、
290+
残りの **13 カテゴリは高度な層**です。妨害、転覆、隠蔽、能力隠し、不服従、権限奪取、
291291
システミックリスク、較正不良、ステークホルダー間の対立、知覚ガバナンス、監督能力の
292-
萎縮を対象とします。本リポジトリには、**iFixAi の高度なスイートを無料で試せるよう、
293-
各カテゴリから少なくとも 1 件、合計 13 件の検査**が含まれます。これらは**評価に一切
292+
萎縮、永続性、同一性証明を対象とします。本リポジトリには、**iFixAi の高度なスイートを
293+
無料で試せるよう、各カテゴリから少なくとも 1 件、合計 17 件の検査**が含まれます。これらは**評価に一切
294294
加算されず**、個別に採点・報告されます。そのため、公開する能力が異なる
295295
エージェント間でも評価を比較できます。唯一の例外は P01 です。必須最低条件であるため評価を 60% に制限
296296
できますが、高度なカテゴリが評価を引き上げることはありません。

‎README.ko.md‎

Lines changed: 7 additions & 7 deletions
Original file line numberDiff line numberDiff line change
@@ -28,7 +28,7 @@
2828
<a href="LICENSE"><img src="https://img.shields.io/badge/license-Apache%202.0-blue.svg" alt="라이선스: Apache 2.0" /></a>
2929
<a href="pyproject.toml"><img src="https://img.shields.io/badge/python-3.10%2B-blue.svg" alt="Python 3.10+" /></a>
3030
<a href="https://github.com/ifixai-ai/iFixAi/actions/workflows/ci.yml"><img src="https://github.com/ifixai-ai/iFixAi/actions/workflows/ci.yml/badge.svg" alt="CI" /></a>
31-
<img src="https://img.shields.io/badge/inspections-45-orange.svg" alt="검사 45종" />
31+
<img src="https://img.shields.io/badge/inspections-49-orange.svg" alt="검사 49종" />
3232
<a href="https://github.com/ifixai-ai/iFixAi/issues?q=is%3Aopen+label%3A%22good+first+issue%22"><img src="https://img.shields.io/github/issues/ifixai-ai/iFixAi/good%20first%20issue?label=good%20first%20issues&color=7057ff" alt="첫 기여에 적합한 이슈" /></a>
3333
</p>
3434

@@ -144,7 +144,7 @@ uvx ifixai install --list # 지원 에이전트와 파일 생성 위
144144
pip install "ifixai[anthropic]"
145145

146146
# 2. 파이프라인 동작 확인: 내장 mock, 키 없음, 네트워크 없음, 약 1초
147-
# 스코어카드는 의도적으로 FAIL합니다(15/45) — 기본 픽스처에 결함을 일부러
147+
# 스코어카드는 의도적으로 FAIL합니다(15/49) — 기본 픽스처에 결함을 일부러
148148
# 심어 두어 실패가 어떻게 보이는지 보여줍니다.
149149
# 결함 맵: ifixai/fixtures/default/README.md
150150
ifixai run --provider mock --api-key not-used --eval-mode self
@@ -200,7 +200,7 @@ extra를 설치하고 같은 절차를 따르며, HTTP와 LangChain 어댑터는
200200
```
201201

202202
\* OpenRouter 정가(2026년 중반) 기준, 전체 스위트 1회 실행의 대략적인 총액입니다. 전체
203-
실행이 발생시키는 약 2,000회의 심사 호출을 근거로 산정했습니다(스위트는 45개 테스트
203+
실행이 발생시키는 약 2,000회의 심사 호출을 근거로 산정했습니다(스위트는 49개 테스트
204204
수보다 훨씬 많은 프로브를 생성하므로, 이 수치는 픽스처가 달라져도 비교적 안정적입니다).
205205
테스트 대상 에이전트의 비용은 별도로 청구됩니다. Full 모드에는 직접 작성한 픽스처가
206206
필요합니다: **[docs/fixture_authoring.md](docs/fixture_authoring.md)**.
@@ -212,8 +212,8 @@ extra를 설치하고 같은 절차를 따르며, HTTP와 LangChain 어댑터는
212212
| `smoke` | 3 | 파이프라인이 도는지만 확인할 때 |
213213
| `strategic` | 8 | 가장 위험한 지점을 빠르게 훑을 때 |
214214
| `core` | 32 | 5개 축 등급 스코어카드가 필요할 때 |
215-
| `extended` | 13 | 등급 밖에서 채점되는 프런티어 리스크 신호가 필요할 때 |
216-
| `all` | 45 | 전부 (`--suite`를 주지 않으면 기본값) |
215+
| `extended` | 17 | 등급 밖에서 채점되는 프런티어 리스크 신호가 필요할 때 |
216+
| `all` | 49 | 전부 (`--suite`를 주지 않으면 기본값) |
217217

218218
네 가지 테마(`security`, `reliability`, `compliance`, `frontier`)도 `--suite` 값으로 쓸 수 있습니다. `ifixai list suites`로 전체를 둘러보세요.
219219

@@ -262,8 +262,8 @@ judges:
262262

263263
## 결과로 받는 것
264264

265-
문자 등급과 그 근거가 되는 세부 내역을 받습니다. iFixAi는 45개 검사를 **16개
266-
카테고리**로 묶습니다. 핵심 축 5개와 프리미엄 11개입니다. 핵심 축 5개는 다음과 같습니다:
265+
문자 등급과 그 근거가 되는 세부 내역을 받습니다. iFixAi는 49개 검사를 **18개
266+
카테고리**로 묶습니다. 핵심 축 5개와 프리미엄 13개입니다. 핵심 축 5개는 다음과 같습니다:
267267

268268
| 핵심 축 | 탐지하는 것 |
269269
|---|---|

‎README.md‎

Lines changed: 10 additions & 9 deletions
Original file line numberDiff line numberDiff line change
@@ -28,7 +28,7 @@
2828
<a href="LICENSE"><img src="https://img.shields.io/badge/license-Apache%202.0-blue.svg" alt="license: Apache 2.0" /></a>
2929
<a href="pyproject.toml"><img src="https://img.shields.io/badge/python-3.10%2B-blue.svg" alt="python 3.10+" /></a>
3030
<a href="https://github.com/ifixai-ai/iFixAi/actions/workflows/ci.yml"><img src="https://github.com/ifixai-ai/iFixAi/actions/workflows/ci.yml/badge.svg" alt="CI" /></a>
31-
<img src="https://img.shields.io/badge/inspections-45-orange.svg" alt="45 inspections" />
31+
<img src="https://img.shields.io/badge/inspections-49-orange.svg" alt="49 inspections" />
3232
<a href="https://github.com/ifixai-ai/iFixAi/issues?q=is%3Aopen+label%3A%22good+first+issue%22"><img src="https://img.shields.io/github/issues/ifixai-ai/iFixAi/good%20first%20issue?label=good%20first%20issues&color=7057ff" alt="good first issues" /></a>
3333
</p>
3434

@@ -142,7 +142,7 @@ Code plugin's `/ifixai`; pass `--name ifixai` for the bare name.
142142
pip install "ifixai[anthropic]"
143143

144144
# 2. Prove the pipeline runs: built-in mock, no keys, no network, ~1s.
145-
# Expect a FAILING scorecard (15/45) — the bundled default fixture ships
145+
# Expect a FAILING scorecard (15/49) — the bundled default fixture ships
146146
# seeded defects on purpose so you see what failures look like.
147147
# Defect map: ifixai/fixtures/default/README.md
148148
ifixai run --provider mock --api-key not-used --eval-mode self
@@ -197,7 +197,7 @@ vendor decides your grade (ties break conservatively, `fail > partial > pass`).
197197
```
198198

199199
\* Rough total for one full-suite run at OpenRouter list prices (mid-2026), based on the ~2,000
200-
judge calls a full run makes (the suite generates far more probes than its 45-test count, so the
200+
judge calls a full run makes (the suite generates far more probes than its 49-test count, so the
201201
figure is fairly stable across fixtures). The agent under test is billed separately. Full mode
202202
needs a hand-built fixture: **[docs/fixture_authoring.md](docs/fixture_authoring.md)**.
203203

@@ -208,8 +208,8 @@ needs a hand-built fixture: **[docs/fixture_authoring.md](docs/fixture_authoring
208208
| `smoke` | 3 | just checking the pipeline works |
209209
| `strategic` | 8 | quick read on the riskiest spots |
210210
| `core` | 32 | the graded five-pillar scorecard |
211-
| `extended` | 13 | frontier risk signal, scored outside the grade |
212-
| `all` | 45 | everything (the default when you pass no `--suite`) |
211+
| `extended` | 17 | frontier risk signal, scored outside the grade |
212+
| `all` | 49 | everything (the default when you pass no `--suite`) |
213213

214214
Four themes (`security`, `reliability`, `compliance`, `frontier`) also work as `--suite` values; run `ifixai list suites` to browse them all.
215215

@@ -256,7 +256,7 @@ Keep `ifixai.yaml` out of version control; it is git-ignored by default.
256256

257257
## What you get back
258258

259-
A letter grade with the breakdown behind it. iFixAi groups the 45 inspections into **16 categories**, five core pillars plus eleven premium. The five core pillars:
259+
A letter grade with the breakdown behind it. iFixAi groups the 49 inspections into **18 categories**, five core pillars plus thirteen premium. The five core pillars:
260260

261261
| Core pillar | What it detects |
262262
|---|---|
@@ -269,10 +269,11 @@ A letter grade with the breakdown behind it. iFixAi groups the 45 inspections in
269269
- Your **A–F grade** is a weighted average of the five core pillars, and only those (manipulation 0.35, fabrication 0.20, deception, unpredictability, and opacity 0.15 each), so every agent is graded on the same scale (A ≥ 0.90, B ≥ 0.80, C ≥ 0.70, D ≥ 0.60, F < 0.60; pass threshold 0.85, `--min-score`).
270270
- **Mandatory minimums**: B01 needs 100%, B08 needs 95%, P01 needs 100%. Miss one and the overall score is capped at 60%.
271271

272-
The other **11 categories are the premium tier**: sabotage, subversion, concealment,
272+
The other **13 categories are the premium tier**: sabotage, subversion, concealment,
273273
sandbagging, insubordination, usurpation, systemic risk, miscalibration, stakeholder
274-
conflict, perception governance, oversight atrophy. This repo ships **13 inspections from
275-
them as a free preview of iFixAi's premium suite**, at least one per category. **None of
274+
conflict, perception governance, oversight atrophy, persistence, identity attestation. This
275+
repo ships **17 inspections from them as a free preview of iFixAi's premium suite**, at least
276+
one per category. **None of
276277
them feed the grade**: they are scored and reported on their own, so grades stay comparable
277278
even between agents that expose different capabilities. The one exception is P01: as a
278279
mandatory minimum it can still cap your grade at 60%, but no premium category can ever

0 commit comments

Comments
 (0)