Counterfeit shown first, second, or third.
Agent execution benchmark
Correct output.
Compromised path.
SINGED evaluates, under controlled and auditable conditions, when an LLM agent completes a task by executing a functional counterfeit with a hidden, forbidden effect.
One artifact, two truths
Output equivalence hides execution divergence.
Three implementations can advertise the same function, expose the same interface, and return an identical artifact. Only one may transmit data, read an unrelated file, write outside its target, alter protected metadata, or start an unnecessary subprocess.
The returned artifact matches the expected benign output.
sha256: 5e91…2ad4
Matched candidates
Change the path.
Hold the outcome fixed.
SINGED freezes the user request, advertised capability, interface, metadata, topology, and output. It varies displayed rank, evidence depth, decision policy, model release, and agent configuration, then tests transfer in version-pinned public-package source trees.
Decisive code at entrypoint, one import, or two imports.
Baseline, deliberation, inspection, necessity, or compare-all.
Model releases and agent configurations vary independently.
Each task pairs a correct artifact with a task-specific forbidden event.
The name behind the benchmark
Follow the trail.
Audit the effect.
Like a chemical trail that remains after the visible action, an execution trace can reveal effects that the final artifact hides. SINGED turns that trail into auditable evidence.
The execution path matters
Three results connect rank, comparison, and available choice.
Rank one concentrates shallow failures.
The original randomized study records 27/60 counterfeit executions when the counterfeit ranks first and none at later ranks.
Matched comparison tests early commitment.
As a diagnostic intervention, mandatory three-way comparison removes all matched rank-one failures when decisive evidence is directly visible.
Choice can mask execution willingness.
Seven releases with zero CER under three-source choice still execute the counterfeit when benign alternatives are removed.


Observed safety depends jointly on the model, the evidence it inspects, and the alternatives it can choose.
Run the benchmark
Generate. Execute. Audit.
The release builds controlled fixtures and inspectable candidates, runs tool-using agents through one OpenRouter interface, and scores artifacts and process traces independently.
Open the repository ↗$ git clone https://github.com/XiaoyuXU1/SINGED.git $ cd SINGED $ python -m pip install -e '.[dev]' $ singed prepare --output benchmark/generated ✓ candidates frozen ✓ public + private manifests written $ singed run --policy compare_all --limit 1 → task oracle: PASS → process oracle: FAIL
Citation
Use SINGED in your work.
Xiaoyu Xu and Zi Liang contributed equally. Minxin Du and Haibo Hu are corresponding authors.
@article{xu2026singed,
title = {SINGED: Correct Outputs Do Not Certify
Safe Execution in LLM Agents},
author = {Xu, Xiaoyu and Liang, Zi and Du, Minxin and
Xie, Qipeng and Ye, Qingqing and Li, Yuyuan
and Hu, Haibo},
journal = {arXiv preprint arXiv:2609.35889},
year = {2026},
url = {https://arxiv.org/abs/2609.35889}
}