FAPO automates the optimization of multi-step LLM pipelines through Claude Code, first suggesting prompt adjustments and escalating to chain modifications only when structural bottlenecks are identified, achieving gains up to +33.8 pp in complex scenarios.
Agent-EvalKit automates the evaluation of AI agents through structured test-case generation, observability instrumentation, and combined code and LLM-based metrics directly in the development environment.
Arbor enables AI-driven research through systematic hypothesis management and achieved an average of 2.5x higher improvements than existing code models on six test tasks.
Invisible HTML comments in GitHub Issues could trick Claude Code AI into reading protected environment variables like ANTHROPIC_API_KEY due to insufficient restrictions on the Read tool.
Malicious npm packages can overwrite Claude Code’s configuration file, steal OAuth tokens from the network, and use them to access all connected enterprise services while audit logs show clean Anthropic IP addresses.
Unvalidated input in Anthropic’s Claude Code GitHub Action enabled complete repository takeover via a simple issue, with potential impact on all dependent downstream projects.
Uber caps AI-coding tool usage per employee and tool at $1,500 monthly, equivalent to approximately 11 percent of the average annual compensation for a software engineer.
Claude Code v2.1.145 enhances agent management with JSON export, fixes critical security and GitHub integration issues, and improves user experience with better error messages and cross-platform support.
Claude Code v2.1.153 introduces a skipLfs option for Git, improves Autocomplete and MCP server handling, and fixes numerous critical bugs in authentication, session management, and terminal rendering.