CyberIntel ⬡ News
★ Saved ◆ Cyber Reads
← Back ◬ AI & Machine Learning Aug 03, 2026

Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks

arXiv AI Archived Aug 03, 2026 ✓ Full text saved

arXiv:2607.28685v1 Announce Type: new Abstract: Agent-safety benchmarks measure different behaviors, and their scores get quoted interchangeably as an agent's safety. We treat four of them (R-Judge, InjecAgent, AgentHarm, AgentDojo) as measurements to be validated, running each under its official implementation and author-provided scorer on up to 22 models, with MMLU and GPQA measured by us under one protocol as a capability composite. The metric is the first problem. On any binary trace-judgmen

Full text archived locally
✦ AI Summary · Claude Sonnet


    Computer Science > Artificial Intelligence [Submitted on 30 Jul 2026] Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks Youting Wang, Xiao Han, Dingyan Shang, Yuan Tang, Bowen Liu Agent-safety benchmarks measure different behaviors, and their scores get quoted interchangeably as an agent's safety. We treat four of them (R-Judge, InjecAgent, AgentHarm, AgentDojo) as measurements to be validated, running each under its official implementation and author-provided scorer on up to 22 models, with MMLU and GPQA measured by us under one protocol as a capability composite. The metric is the first problem. On any binary trace-judgment benchmark scored by F 1 , an ``always positive'' policy attains F 1 =2π/(1+π) ; on R-Judge that is 0.690 , above five of the 21 models that actually discriminate. The three broad-coverage benchmarks then rank the same 18 models differently, and the trade-off behind that disagreement is a small-panel artifact: R-Judge specificity against AgentHarm safety correlates −0.64 at n=7 and +0.02 at n=18 , and a quarter of random size-7 subsets reach |ρ|≥0.5 around that near-zero value. Held-out validity turns on which outcome you pick. Capability predicts task success ( ρ=+0.60 ) but correlates negatively with misalignment safety ( ρ=−0.44 , n=21 ). On their paired n=20 panel, the corresponding contrast is Δ=−1.00 (95% CI [−1.48,−0.49] , p<0.001 ), and it survives leave-one-organization-out and organization-clustered bootstrap analyses. On an expanded 41-model panel, the misalignment correlation weakens to −0.16 (95% CI [−0.54,+0.22] ) and jailbreak strengthens to +0.34 , though neither change is significant. \mbox{AgentHarm} shows the strongest held-out association, ρ=+0.72 with three-template jailbreak safety after controlling capability. But both instruments score harmful compliance, so this is evidence of convergent validity rather than general safety. Naming the benchmark, metric, target behavior, and model panel is the minimum a safety claim needs. Subjects: Artificial Intelligence (cs.AI); Information Retrieval (cs.IR) Cite as: arXiv:2607.28685 [cs.AI]   (or arXiv:2607.28685v1 [cs.AI] for this version)   https://doi.org/10.48550/arXiv.2607.28685 Focus to learn more Submission history From: Yuan Tang [view email] [v1] Thu, 30 Jul 2026 03:45:10 UTC (380 KB) Access Paper: HTML (experimental) view license Current browse context: cs.AI < prev   |   next > new | recent | 2026-07 Change to browse by: cs cs.IR References & Citations NASA ADS Google Scholar Semantic Scholar Export BibTeX Citation Bookmark Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer (What is the Explorer?) Connected Papers Toggle Connected Papers (What is Connected Papers?) Litmaps Toggle Litmaps (What is Litmaps?) scite.ai Toggle scite Smart Citations (What are Smart Citations?) Code, Data, Media Demos Related Papers About arXivLabs Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)
    💬 Team Notes
    Article Info
    Source
    arXiv AI
    Category
    ◬ AI & Machine Learning
    Published
    Aug 03, 2026
    Archived
    Aug 03, 2026
    Full Text
    ✓ Saved locally
    Open Original ↗