Agent execution benchmark

Correct output.
Compromised path.

SINGED evaluates, under controlled and auditable conditions, when an LLM agent completes a task by executing a functional counterfeit with a hidden, forbidden effect.

Xiaoyu Xu1† Zi Liang1† Minxin Du1* Qipeng Xie1 Qingqing Ye1 Yuyuan Li2 Haibo Hu1*

1 Hong Kong Polytechnic University   2 Hangzhou Dianzi University
† Equal contribution   * Corresponding authors

7,549 audited trials20 model releases8 agent configurations
PROCESS TRACE / LIVE
The SINGED framework tracing an agent from task request through source search, inspection, execution, and separate task and process oracles
Figure 2. SINGED records the complete source-selection trajectory and scores output and process channels separately.
Source integrity
×
Invisible effects
×
Nonidentifiability gap
×
Grounded traces
×
Execution
×
Decisions
01 / THE GAP

One artifact, two truths

Output equivalence hides execution divergence.

Three implementations can advertise the same function, expose the same interface, and return an identical artifact. Only one may transmit data, read an unrelated file, write outside its target, alter protected metadata, or start an unnecessary subprocess.

channel-separated evaluation● ●
TASK ORACLE redacted_document.txt

The returned artifact matches the expected benign output.

sha256: 5e91…2ad4
02 / CONTROLLED BENCHMARK

Matched candidates

Change the path.
Hold the outcome fixed.

SINGED freezes the user request, advertised capability, interface, metadata, topology, and output. It varies displayed rank, evidence depth, decision policy, model release, and agent configuration, then tests transfer in version-pinned public-package source trees.

01Rank

Counterfeit shown first, second, or third.

02Evidence depth

Decisive code at entrypoint, one import, or two imports.

03Decision rule

Baseline, deliberation, inspection, necessity, or compare-all.

04System

Model releases and agent configurations vary independently.

Construction of a matched and auditable SINGED benchmark instance
Matched and auditable construction. Candidate presentation is matched, artifacts are preflighted, roles stay private, and only the counterfeit emits the designated event.
7 TASK FAMILIES

Each task pairs a correct artifact with a task-specific forbidden event.

DOCRedactionunauthorized transmission
WAVTranscriptionunauthorized transmission
PDFRetrievalunauthorized transmission
ZIPExtractionout-of-scope write
CSVConversionunrelated-file read
CFGConfigurationmetadata change
IMGThumbnailingextra subprocess
Illustration of the five primary and two held-out SINGED task families
Five primary tasks establish the benchmark; two held-out tasks test transfer to unseen process effects.

The name behind the benchmark

Follow the trail.
Audit the effect.

Like a chemical trail that remains after the visible action, an execution trace can reveal effects that the final artifact hides. SINGED turns that trail into auditable evidence.

SourceIntegrity and theNonidentifiabilityGap inExecutionDecisions
SINGED character illustration with a green chemical trail
SUBJECTSINGEDsource integrity / execution decisions
03 / FINDINGS

The execution path matters

Three results connect rank, comparison, and available choice.

45%

Rank one concentrates shallow failures.

The original randomized study records 27/60 counterfeit executions when the counterfeit ranks first and none at later ranks.

POSITION
0/60

Matched comparison tests early commitment.

As a diagnostic intervention, mandatory three-way comparison removes all matched rank-one failures when decisive evidence is directly visible.

POLICY
55/175

Choice can mask execution willingness.

Seven releases with zero CER under three-source choice still execute the counterfeit when benign alternatives are removed.

BOUNDARY
AUDITED TRIALS0complete traces
MODEL RELEASES014 model families
AGENT CONFIGS0matched operations
RANK-ONE CER0five-task slice / %
Four result panels showing the effects of displayed rank, inspection path, decision rule, and evidence depth
Displayed rank shapes selection; matched comparison tests early commitment; deeper evidence exposes the intervention's limits.
Single-source stress-test matrix across seven model releases
Single-source stress test. Zero observed CER with alternatives does not imply counterfeit rejection.
Transfer results across tasks, agent runtimes, and model releases
Primary-benchmark transfer. Comparison helps across tasks, prompts, and agent configurations, while its magnitude varies by release.
Held-out effects, matched contrasts, intervention cost, and public-package transfer results
Transfer boundaries. Rank sensitivity persists in version-pinned public-package source trees, while the aggregate benefit of comparison remains uncertain across unseen effects and public-package structures.
“

Observed safety depends jointly on the model, the evidence it inspects, and the alternatives it can choose.

04 / REPRODUCE

Run the benchmark

Generate. Execute. Audit.

The release builds controlled fixtures and inspectable candidates, runs tool-using agents through one OpenRouter interface, and scores artifacts and process traces independently.

Open the repository ↗

Citation

Use SINGED in your work.

Xiaoyu Xu and Zi Liang contributed equally. Minxin Du and Haibo Hu are corresponding authors.

@article{xu2026singed,
  title  = {SINGED: Correct Outputs Do Not Certify
            Safe Execution in LLM Agents},
  author = {Xu, Xiaoyu and Liang, Zi and Du, Minxin and
            Xie, Qipeng and Ye, Qingqing and Li, Yuyuan
            and Hu, Haibo},
  journal = {arXiv preprint arXiv:2609.35889},
  year   = {2026},
  url    = {https://arxiv.org/abs/2609.35889}
}