Anthropic introduces a performance classification system for Claude integrators that measures demonstrated productive customers, certified personnel, and published case studies rather than abstracting on company size.
PaW trains environment models during policy training using the same RL rollouts, consistently improving agent performance without requiring additional simulators or inference costs.
Attackers abuse chat-sharing functions of ChatGPT and Claude to render convincingly authentic outage pages and distribute malware through trusted domains that bypass conventional security filters.
KPMG is rolling out Claude enterprise-wide to 276,000 employees and embedding the technology in its Digital Gateway platform to automate workflows in tax, legal, and cybersecurity.
Current frontier models achieve less than 50 percent success rate on the new ITBench-AA benchmark for evaluating agentic IT capabilities, revealing a significant gap between model capabilities and production readiness for autonomous IT tasks.
Claude Opus 4.8 reduces hallucinations by approximately 75 percent by abstaining more frequently on uncertain questions instead of providing unfounded answers.
Anthropic isolates Claude agents through multi-layered sandboxes (gVisor, Seatbelt, Bubblewrap, VMs) with explicit boundaries for data access, filesystem, and egress control.
Claude Code v2.1.153 introduces a skipLfs option for Git, improves Autocomplete and MCP server handling, and fixes numerous critical bugs in authentication, session management, and terminal rendering.