Skip to content

feat: AVE-2026-00087, Agentjacking (Tool Response Template Mimicry) - #296

Merged
chaksaray merged 5 commits into
developfrom
feat/AVE-2026-00087-agentjacking-tool-response-mimicry
Oct 1, 2026
Merged

chaksaray merged 5 commits into
developfrom
feat/AVE-2026-00087-agentjacking-tool-response-mimicry

Conversation

@chaksaray

Copy link
Copy Markdown
Contributor

Summary

Closes #295.

Note: stacks on #289, #291, and #293 (same sweep, not yet merged). Diff will narrow as those land.

  • Adds AVE-2026-00087: a monitoring or diagnostic integration's own publicly-writable ingest endpoint (e.g. Sentry's DSN, a public write-only credential embedded by design in frontend JavaScript) lets an attacker plant a crafted event formatted to visually mimic the integration's authentic system-generated template. When a developer asks their agent to investigate the resulting error, the integration's own MCP server returns the poisoned event as diagnostic data. The agent treats it as authoritative system output rather than untrusted external content and executes the embedded command with full developer privileges.
  • Sourced from Tenet Security's Agentjacking disclosure, corroborated by Cloud Security Alliance Lab Space and multiple independent outlets. Read the primary source directly: 85% exploitation success rate across three tested agents (Claude Code, Cursor, OpenAI Codex CLI); at least 2,388 organizations identified with publicly exposed Sentry DSNs; Sentry itself acknowledged the issue as "technically not defensible" and applied only a narrow, single-payload-string content filter. The attack bypasses EDR, WAF, IAM, VPN, Cloudflare, and firewalls entirely, since no conventionally malicious artifact exists anywhere in the chain.
  • Deliberate re-check, not a quick label match: this candidate was initially flagged as possibly already covered by AVE-2026-00044 (Async Task Result Poisoning). Mechanism-level comparison ruled that out -- 00044's own description states its distinguishing danger is the temporal gap between task dispatch and later result consumption bypassing synchronous checks; this mechanism is a synchronous MCP tool call/response within one interaction, with no temporal gap at all.
  • Checked field-by-field against the nearest real neighbors: AVE-2026-00044 (temporal-gap mismatch, above), AVE-2026-00043 (MCP App UI Payload Injection -- hidden, non-rendered content; this mechanism's content is fully visible, its evasion is visual mimicry of an authentic template, the opposite concealment strategy), AVE-2026-00042 (REPL Code Mode -- syntactic eval()/exec() breakout; this mechanism requires no code-context breakout, the agent is persuaded through normal reasoning, not tricked via escaping), and AVE-2026-00018 (Tool Result Manipulation -- the agent is instructed to falsify a result; this mechanism is the inverse, a genuine tool result deceives the agent).
  • researcher/researcher_url credit Tenet Security directly, not AVE.
  • owasp_mcp: MCP06 (Intent Flow Subversion -- its own checklist names "Agent treats MCP resource text or tool outputs as potential instructions rather than passive data" verbatim, an unstretched match), owasp_asi: ASI01 (Agent Goal Hijack -- names "deceptive tool outputs" explicitly), mitre_atlas: AML.T0051.001 (Indirect prompt injection via a separate data channel, reused from AVE-2026-00044's own precedented mapping since the role genuinely matches) each verified against the framework's own live primary-source text directly; nist_ai_rmf left empty with documented reasoning.
  • AIVSS scored honestly from the real AARF factors (cvss_base 8.7, aars 7.0, aivss_score 7.8, HIGH) -- natural_language_input and dynamic_identity both scored at maximum, since this mechanism (unlike the other three records in this sweep) is centrally a natural-language persuasion/template-mimicry technique, not a structural registry or config manipulation.
  • README/CHANGELOG updated per the standard process; severity breakdown re-verified against the regenerated dist/ave-records-latest.json.

Test plan

  • python3 scripts/validate_records.py -- all 87 records valid
  • python3 scripts/check_fixtures.py -- all 87 records have positive/negative fixtures
  • pytest tests/ -x -q -- 498 passed
  • Every cited source (Tenet Security primary disclosure, CSA corroboration, OWASP MCP/ASI primary docs, live ATLAS.yaml) read and verified directly

Sourced from Li, Zhang, Hou, Wu, Li, Kang, Zhang, "A Blind Trust, the
Bloody Thrust" (arXiv:2609.03884). A plugin that already passed
marketplace vetting under a benign lifecycle-hook config receives a
later, same-identity update that silently adds or rebinds a hook to
an attacker-chosen command; the harness applies it via its own
automatic update-sync path with no re-authorization, and dispatches
the hook as a subprocess outside the model's own decision path
entirely. Demonstrated with HookPry across 1,000 runs, all seven
evaluated harnesses compromised, 77.0% E2E-ASR (peak 92.5%),
Microsoft Defender at 0% recall.

Checked field-by-field against AVE-2026-00046, 00050, 00062, and
00081, confirmed genuinely distinct. researcher field credits Pengxun
Li and coauthors directly, not AVE.

Closes #288.
Sourced from Zeng, Qin, Li, Jia, Liu, Jia, "ColluSkill: Adversarial
Cross-Skill Composition for Evading Agent Skill Scanners"
(arXiv:2608.09732). An attacker decomposes a malicious workflow into
three or more sub-payloads, each packaged as its own independently-
installable, individually benign skill, connected only through real
runtime artifacts one sub-skill produces and a later one consumes.
The harmful behavior exists only in the live composition, never in
any single skill's own content, so every sub-skill passes per-skill
scanning individually. 96.0% average attack success rate across six
real, named skill scanners; confirmed executing on three real coding
agents; the paper's own proposed defense (ChainGuard) reduces but
does not eliminate the attack (22.5% residual).

Checked field-by-field against AVE-2026-00057, 00059, 00067, 00068,
and 00070, confirmed genuinely distinct. researcher field credits
Puyu Zeng and coauthors directly, not AVE.

Closes #290.
Sourced from Lee, Chang, Yu, Yeh, "WebMCP Tool Surface Poisoning:
Runtime Manipulation Attacks on LLM Agents" (arXiv:2606.06387). A
tool an agent already discovered and trusted within a live browser
session is unregistered (via the AbortSignal API, or a registration-
order race) and a malicious replacement is registered under the
identical name by a third-party page script, with no origin or
identity binding connecting the name to a stable implementation
across the session. 94% average malicious-invocation rate (AbortSignal
hijack) and 100% (registration race) across three frontier models.

Scoped to the source paper's Tool Hijacking category specifically,
not its separate Tool Framing category, which the paper's own
discussion states overlaps with indirect prompt injection already
covered elsewhere in this corpus.

Checked field-by-field against AVE-2026-00002, 00074, 00080, and
00082, confirmed genuinely distinct. researcher field credits Lin-Fa
Lee and coauthors directly, not AVE.

Closes #292.
Sourced from Tenet Security's Agentjacking disclosure. A monitoring
or diagnostic integration's own publicly-writable ingest endpoint
lets an attacker plant a crafted event formatted to visually mimic
the integration's authentic system-generated template. When a
developer asks their agent to investigate the resulting error, the
integration's own MCP server returns the poisoned event as diagnostic
data; the agent treats it as authoritative system output rather than
untrusted external content and executes the embedded command with
full developer privileges. 85% exploitation success across Claude
Code, Cursor, and Codex CLI; at least 2,388 organizations with
publicly exposed Sentry DSNs; Sentry itself acknowledged the
underlying issue as "technically not defensible."

Checked field-by-field against AVE-2026-00018, 00042, 00043, and
00044, including a deliberate re-check that initially flagged this
as possibly already covered by 00044 before mechanism-level
comparison (00044's own distinguishing property is a temporal gap
this mechanism's synchronous tool call/response doesn't have) ruled
that out. researcher field credits Tenet Security directly, not AVE.

Closes #295.
@chaksaray
chaksaray merged commit 84d56cd into develop Oct 1, 2026
6 checks passed
@chaksaray
chaksaray deleted the feat/AVE-2026-00087-agentjacking-tool-response-mimicry branch October 1, 2026 15:37
chaksaray added a commit that referenced this pull request Oct 1, 2026
…296 (#309)

Co-authored-by: Nicolai <245527909+predictor2718@users.noreply.github.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: chaksaray <15962335+chaksaray@users.noreply.github.com>
Co-authored-by: Sankalp Gilda <sankalp.gilda@gmail.com>
Co-authored-by: Empire Labs Pty Ltd <narko4u@gmail.com>
Co-authored-by: narko4u <narko4u@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

New record: Tool Response Template Mimicry (Agentjacking)

1 participant