Core implementation of EVO-Detect, the dual-flow backdoor detection method described in Beyond Static Cues: Unified Dynamic Detection of MLLM Backdoors via Dual-Flow Trajectory Analysis.
This repository intentionally contains only the online method implementation:
- Attention Flow extraction over early decoding steps;
- Logits Flow extraction with full, text-only, and image-only inputs;
- trajectory aggregation and sample feature construction;
- control-aware calibration, scoring, and thresholding.
Dataset preparation, poisoning, training, benchmark adapters, model-specific data loaders, baselines, plotting, experiment orchestration, and paper results are intentionally excluded. This is a compact core-code release, not a full artifact for reproducing every table in the paper.
pip install -r requirements.txt
pip install -e . --no-depsThe implementation is extracted from the project's online evaluation code and retains its original behavior. It supports the model-loading paths present in that implementation, including Qwen-VL-style and InternVL-style Hugging Face checkpoints. No additional compatibility layer is provided.
src/evo_detect/
├── flows.py # model loading and Attention/Logits Flow extraction
├── scoring.py # calibration, scoring, and decision utilities
└── __init__.py # public API
evo_detect.flows exposes:
load_modelextract_attention_flowextract_logits_flowgenerate_outputaggregate_flowsbuild_sample_feature
evo_detect.scoring exposes the reference trajectory, metric, calibration,
dual-trigger scoring, and threshold utilities used by the online pipeline.
The default early-decoding horizon in the paper is five tokens; callers pass
that horizon through max_new_tokens when extracting each flow.
The caller is responsible for loading samples and images, supplying clean reference and matched-control groups, and persisting any scores or filtered datasets. The implementation does not inspect dataset labels or impose a JSON schema.
MIT. See LICENSE.