GPT-Red detects indirect prompt injections with 84 percent success rate — significantly higher than the 13 percent achieved by human experts — and contributes to developing safer model generations.
AI security depends on governance and human accountability, not model size; increased computing power without organizational control only amplifies existing weaknesses.
Anthropic integrated a hidden tracking mechanism in Claude Code to identify unauthorized resellers and model scrapers — but has already announced its removal.
78 percent of enterprises are already experiencing AI security incidents, while formal governance programmes and central transparency over AI systems remain in the minority.
Security incidents in the AI environment are widespread, but lack of central transparency and governance prevent organizations from fully controlling risks.
Anthropic identifies J-Space as a central cognitive space in Claude that exhibits five characteristics of human consciousness and can be examined using the Jacobian Lens tool.