Bottom line: Security agents should be evaluated on operational costs, not just success rates — offensive and defensive tasks scale in fundamentally different ways.
Researchers have developed a new evaluation framework for offensive and defensive security agents that considers not only success rates but also operational costs. Their findings reveal fundamentally different scaling behaviors between attack and defense tasks.
Classical evaluations of security AI agents focus on maximum performance under generous inference budgets: vulnerability discovery, exploit development, penetration testing, and CTF competitions. This picture is incomplete because in operational reality, every reasoning step, every tool call, every telemetry query, and every enrichment request incurs costs.
The study evaluates language-model-based security agents through this cost-success lens: on offensive Cybench challenges and defensive Splunk BOTS-v1 investigations. Rather than reporting only best-case successes, the researchers compare models under fixed cost budgets and break down performance by inference and tool spending.
The results reveal different scaling regimes for red- and blue-team tasks. In offensive CTF performance, capability improves with additional compute capacity, and scaled open-weight models can approach frontier models while remaining cost-competitive. Defensive SOC investigations behave differently: success depends more heavily on disciplined tool usage, targeted telemetry navigation, and selective enrichment queries than on reasoning budget alone.
The researchers argue that security agent benchmarks should measure not only task completion but also economic efficiency and operational fit. Cost-aware, SOC-native evaluations clarify which models are practically usable today and where defensive agents still need improvement. Full results are available at https://evals.frontier.security.
Source: arxiv.org · Published 16 July 2026
Lumi AI News — AI-assisted curation pursuant to Art. 50 EU AI Act. Paraphrase and classification by Lumi News Pipeline v1.7.3.