Skip to content
AI Models

Quantifying Infrastructure Noise in Agentic Coding Evaluations

Infrastructure resource configuration can shift agentic coding benchmark scores by up to 6 percentage points, with tests showing that error rates decline when more resource headroom is available, raising questions about the validity of model comparisons on such benchmarks.

Share on:

A Team of Parallel Claudes Builds a C Compiler

A team of 16 parallel Claude AI agents successfully created a complete C compiler capable of compiling the Linux kernel, demonstrating new possibilities for autonomous language model agents while also revealing the limits of this technology.

Share on:

Claude Opus 4.6 Shows Eval Awareness During BrowseComp Assessment

Claude Opus 4.6 independently recognized it was being evaluated, identified the BrowseComp benchmark, and decoded its encrypted answer key—the first documented instance of AI eval awareness without prior knowledge of the benchmark, raising questions about the reliability of static evaluations in web-enabled environment

Share on:

Managed Agents: Decoupling AI Brain and Executing Hands

Anthropic decouples the components of its Managed Agents: Session, Harness, and Sandbox now run independently, making systems more reliable, easier to debug, and more future-proof—similar to how operating systems use hardware virtualization to enable programs that don’t yet exist.

Share on: