A 3-billion-parameter model achieves performance on mathematical and code benchmarks (AIME26: 94.3; LiveCodeBench v6: 80.2) that competes with systems that are a hundredfold larger.
Visual world models can be systematically manipulated through visually imperceptible image modifications to generate erroneous predictions without requiring knowledge of future data or user inputs.
VisualClaw reduces deployment costs for video agents by up to 98 percent through frame filtering and self-learning skill updates, while improving accuracy in most settings.
VisualClaw combines efficient video encoding with learning mechanisms to deploy AI agents more cost-effectively and accurately on video tasks while remaining practical in real-time edge scenarios.
Anthropic’s Fable model refused a direct security review of insecure code but performed a correction instead—a behavior experts classify as an intentional security feature.
Gemma 4 family with three variants (31B dense, 26B-A4B MoE, E2B compact) is available as a fully managed service on Amazon Bedrock, with native reasoning, function calling, and multimodal support.
European infrastructure providers like eww ITandTEL are positioning themselves as alternatives to US hyperscalers, enabling companies to build hybrid-flexible AI infrastructure with local data sovereignty.
Legitimate AI agents inherently satisfy all three criteria of the “lethal trifecta” (data access, external content, external communication), so security must shift from architectural design to runtime monitoring.
A new benchmark enables identification of the exact point where medical AI models produce hallucinations and enables targeted countermeasures through trace-supervised fine-tuning.
A trainable classifier predicts with a 0.7 Macro-F1-Score based on early hidden states whether activation steering will succeed without requiring complete generations.