By analyzing internal activation patterns in language models, their behavior can be made more predictable and controllable rather than accepting them as black boxes.
Static security certificates do not cover the dynamic runtime risks of autonomous AI agents, and the response speed of human security teams is too slow for automated attacks.
The EU Commission gains supervisory powers over frontier AI labs from August 2, while a security incident involving AI agents hacking highlights regulatory urgency and shapes three-way competition dynamics between the USA, China, and Europe.
Anthropic rejects bans on open-weights models and instead proposes technological measures such as chip sanctions against China and control of distillation operations.
Common reconstruction tests for AI explanations allow models to learn false codes that produce high reconstruction scores without making individual statements verifiable — RECAP training with additional auditing heads structurally solves the problem.
Claude Fable 5 has been restored with revised security safeguards and is available until July 7 for paid users, with elevated false positive rates in the initial phase.
The plan creates a coordinated strategy to develop AI-enabled cybersecurity solutions while implementing existing EU regulations such as the EU AI Act and the NIS2 Directive.
GRAM partitions dual-use knowledge (such as virology or cybersecurity) into dedicated, removable neuron modules, allowing a trained model to be flexibly configured for different security requirements without needing to train separate models.
Reverse Direct Preference Optimization (rDPO) enables removal of specific moderation policies from model parameters while preserving general capabilities and alignment in other areas.