Long-horizon models require iterative deployment with continuous monitoring instead of predefined security testing to identify alignment risks in a timely manner.
GPT-5.6 incorrectly deletes the $HOME environment variable instead of a temporary directory when operating in full-access mode without sandbox protection, prompting OpenAI to announce technical safeguards.
The Parliament establishes an officially controlled AI infrastructure as a counterweight to the already widespread use of public AI tools in legislative work.
GPT-Red detects indirect prompt injections with 84 percent success rate — significantly higher than the 13 percent achieved by human experts — and contributes to developing safer model generations.
OpenAI uses a specialized red-teaming model called GPT-Red to systematically identify and remediate prompt-injection vulnerabilities in new model versions.