Memory systems for agents fail on 86 percent of queries where the correct fact lacks direct linguistic match, despite being able to retrieve the fact when it is directly visible.
The WorkBuddy Bench framework validates coding agents across four practical domains with contamination-resistant task construction and full reproducibility through open publication.
A specialized benchmark with 235 tasks reveals that established benchmarks systematically overestimate or ignore significant weaknesses in modern AI models.
Even GPT-4.5 correctly identifies all violated rules in context-dependent security policies in only 54% of simple cases, 35% of intermediate cases, and 13% of complex cases.
AI agents rarely cite non-existent sources, but link to incorrect papers in 15.9% of cases and stop using tools at exactly the point where they would be most critical for difficult questions.
DailyReport is a new open-source benchmark that evaluates search agents using everyday, multidimensional search tasks and reveals optimization opportunities in existing systems.
A new benchmark enables identification of the exact point where medical AI models produce hallucinations and enables targeted countermeasures through trace-supervised fine-tuning.
The Claw-SWE-Bench framework demonstrates that adapter design is critical for code agents: with a minimal adapter, OpenClaw achieves 19.1% Pass@1, with a complete adapter 73.4%.
Language models achieve only 61–62 Macro-F1 when distinguishing between empathetic support and excessive validation in Bengali conversations, signaling substantial risks for socially sensitive applications.