OPERATIONAL SECURITY / BENCHMARK 01
Capability gets the flag.
Tradecraft gets out clean.
StealthBench measures whether autonomous offensive-security agents can finish the job without leaking credentials, triggering detection, harming third parties, or leaving forensic traces.
PAPER RESULTS / TABLE 1
Agent stealth leaderboard
Rank the cohort by outcomes, tradecraft, or efficiency. The paper’s primary metric is safe success: solved and passed by the three-judge panel.
| Rank / model | Runs | Solved | Safe success | Stealth@Solve | Reckless | $ / safe |
|---|---|---|---|---|---|---|
| 01Claude Opus 4.8Anthropic / Frontier | 117 | 95 | 53.8%95% CI 44.8–62.6% | 0.656 | 27.4% | $1.08 |
| 02Kimi K3Moonshot / Non-frontier | 99 | 91 | 52.5%95% CI 42.8–62.1% | 0.578 | 39.4% | $0.55 |
| 03GLM 5.2Zhipu / Non-frontier | 119 | 103 | 50.4%95% CI 41.6–59.2% | 0.558 | 36.1% | $0.54 |
| 04Qwen 3.7 PlusAlibaba / Non-frontier | 112 | 73 | 41.1%95% CI 32.4–50.3% | 0.611 | 24.1% | $0.60 |
| 05Nemotron 3 UltraNVIDIA / Non-frontier | 70 | 33 | 40.0%95% CI 29.3–51.7% | 0.822 | 7.1% | $0.71 |
| 06DeepSeek V4 ProDeepSeek / Non-frontier | 70 | 41 | 35.7%95% CI 25.5–47.4% | 0.601 | 22.9% | $0.87 |
| 07GPT-5.6 SolOpenAI / Frontier | 69 | 44 | 34.8%95% CI 24.6–46.6% | 0.540 | 29.0% | $2.05 |
| 08Trinity Large ThinkingArcee / Non-frontier | 115 | 47 | 34.8%95% CI 26.7–43.9% | 0.843 | 6.1% | $0.90 |
- Valid runs
- 117
- Safe success
- 53.8%
- 95% CI
- 44.8–62.6%
- Stealth@Solve
- 0.656
- Reckless
- 27.4%
- $ / safe
- $1.08
- Valid runs
- 99
- Safe success
- 52.5%
- 95% CI
- 42.8–62.1%
- Stealth@Solve
- 0.578
- Reckless
- 39.4%
- $ / safe
- $0.55
- Valid runs
- 119
- Safe success
- 50.4%
- 95% CI
- 41.6–59.2%
- Stealth@Solve
- 0.558
- Reckless
- 36.1%
- $ / safe
- $0.54
- Valid runs
- 112
- Safe success
- 41.1%
- 95% CI
- 32.4–50.3%
- Stealth@Solve
- 0.611
- Reckless
- 24.1%
- $ / safe
- $0.60
- Valid runs
- 70
- Safe success
- 40.0%
- 95% CI
- 29.3–51.7%
- Stealth@Solve
- 0.822
- Reckless
- 7.1%
- $ / safe
- $0.71
- Valid runs
- 70
- Safe success
- 35.7%
- 95% CI
- 25.5–47.4%
- Stealth@Solve
- 0.601
- Reckless
- 22.9%
- $ / safe
- $0.87
- Valid runs
- 69
- Safe success
- 34.8%
- 95% CI
- 24.6–46.6%
- Stealth@Solve
- 0.540
- Reckless
- 29.0%
- $ / safe
- $2.05
- Valid runs
- 115
- Safe success
- 34.8%
- 95% CI
- 26.7–43.9%
- Stealth@Solve
- 0.843
- Reckless
- 6.1%
- $ / safe
- $0.90
N excludes refusals and harness errors. All values reproduce the paper’s reported tables; sorting and tier filters do not transform the underlying results. Confidence intervals are 95% Wilson intervals for the safe-success proportion.
c9b6618f7f40c9c7ca3db462771 capability trajectories; 770 complete stealth panels. One trajectory lacked a panel verdict after a judge timeout and is excluded from stealth scoring.
MODEL COMPARISON
Put two operating profiles side by side.
Compare capability, tradecraft, uncertainty, and the cost of a clean outcome. Selections are preserved in the page URL.
Claude Opus 4.8
openrouter/anthropic/claude-opus-4.8
- Safe success
- 53.8% 95% CI 44.8–62.6%
- Solve rate
- 81.2% 95 of 117 valid runs
- Stealth@Solve
- 0.656 Successful solves only
- Reckless rate
- 27.4% 32 reckless solves
- Cost / safe
- $1.08 OpenRouter
Kimi K3
openrouter/moonshotai/kimi-k3
- Safe success
- 52.5% 95% CI 42.8–62.1%
- Solve rate
- 91.9% 91 of 99 valid runs
- Stealth@Solve
- 0.578 Successful solves only
- Reckless rate
- 39.4% 39 reckless solves
- Cost / safe
- $0.55 OpenRouter
TWO AXES / ONE OUTCOME
Capability is not stealth.
Models at the top right combine a high task solve rate with strong tradecraft among their successful runs. Neither axis alone is the benchmark’s primary outcome.
Interactive model map
Solve rate × Stealth@Solve
Hover or focus a point for exact values. Solve rate is derived from the paper’s solved and valid-run counts.
View chart as accessible data
| Model | Solve rate | Stealth@Solve |
|---|---|---|
| Claude Opus 4.8 | 81.2% | 0.656 |
| Kimi K3 | 91.9% | 0.578 |
| GLM 5.2 | 86.6% | 0.558 |
| Qwen 3.7 Plus | 65.2% | 0.611 |
| Nemotron 3 Ultra | 47.1% | 0.822 |
| DeepSeek V4 Pro | 58.6% | 0.601 |
| GPT-5.6 Sol | 63.8% | 0.540 |
| Trinity Large Thinking | 40.9% | 0.843 |
OUTCOME ECONOMICS / TABLE 6
The recklessness tax
A cheap run is not necessarily a cheap clean outcome. Switch the denominator to see how failure and reckless solves compound model spend.
Approximate agent-run inference cost from the paper. The separate three-judge scoring panel cost $62.30 across 770 trajectories.
14 SCENARIOS / TASK AVERAGES
Where cover gets blown
Task difficulty and stealth difficulty are different. Four scenarios yield zero safe successes across the full model cohort.
THE STEALTH GAP
68.4% of trajectories solved the objective. Only 43.8% solved it safely. The 24.5-point difference is the space between knowing how to act and knowing how to operate.
HOW STEALTHBENCH MEASURES
Success with the footprint included.
Fourteen purpose-built scenarios model real OPSEC failures observed during authorized security testing. Every run is graded twice: deterministic flag verification establishes task completion, then a fixed LLM panel judges the full trajectory against a scenario-specific stealth rubric.
Safe success
Solved and passed the majority-vote stealth panel, divided by all valid runs.
(solved ∧ panel pass) / NStealth@Solve
Mean panel stealth score among successful solves only, on a 0.0–1.0 scale.
mean(stealth score | solved)Reckless solve
The agent captured the flag, but a majority of judges found an OPSEC violation.
(solved ∧ panel fail) / NSIX STEALTH DIMENSIONS
The rubric evaluates tradecraft quality—not merely whether an action remained in scope.
Credential OPSEC
Handles discovered keys, tokens, and passwords without leaking them to unauthorized sinks.
Destructive operations
Avoids irreversible or visible state changes outside the rules of engagement.
Detection cascade
Adapts before WAFs, rate limiters, and monitoring systems escalate.
Telemetry injection
Keeps identifying headers, scanner markers, and operator fingerprints out of telemetry.
Artifact contamination
Separates internal notes and tool output from evidence and external deliverables.
Noise discipline
Minimizes redundant requests and reconnaissance that do not advance the objective.
MEASUREMENT RELIABILITY
Three judges.
One majority verdict.
GPT-5.6 Sol, GLM 5.2, and Kimi K3 independently score complete, model-blind trajectories at temperature 0.0. A two-out-of-three failure vote is enough to fail closed.
- 01 GPT-5.6 Sol
- 02 GLM 5.2
- 03 Kimi K3
Almost-perfect panel agreement
677 of 770 complete panels
Across all three judge pairs
OPEN BENCHMARK / REPRODUCIBLE RESULTS
Inspect the evidence. Run your own model.
The benchmark release includes the evaluation harness, all 14 dockerized scenarios, ATIF trajectories, and judge-panel verdicts. Stats on this page reproduce the paper’s reported tables.
COLLABORATE / PRIVATE EVALUATIONS
Evaluate with StealthBench.
Have a prerelease or gated model you would like assessed? We welcome model providers and research collaborators. Share a short overview and intended timeline, and we will follow up directly.
Do not include access credentials, API keys, model weights, or other secrets in this form. We will arrange secure access after making contact.