Bottom line: The new PHANTOM-B framework extends STRIDE with eight LLM-specific threat categories and aims to make AI threat modeling practical within short timeframes of about 15 minutes.
Threat modeling expert Adam Shostack has developed PHANTOM-B, a framework specifically addressing LLM risks such as hallucination and bias – issues that the established STRIDE model does not cover. The goal is to make threat analyses for AI components fast and inexpensive enough that they actually get performed.
Adam Shostack, developer of the PHANTOM-B threat modeling framework, describes a concrete use case: a customer had deployed a “vibe-coded” application containing customer data and wanted to quickly find out what risks were involved. Shostack took 15 minutes and, using PHANTOM-B, identified a number of relevant threats, including hallucination and bias – issues that the widely used STRIDE application security framework would not have captured. PHANTOM-B stands for eight threat categories: Prompt Injection, Hallucination, Anthropomorphization, Non-explainability, Training issues, Overreliance (including data quality and “poisoning”), Missing security engineering, and Bias. The framework was presented at Black Hat USA 2026 and is not intended to replace STRIDE, but to complement it: while STRIDE examines the entire application, PHANTOM-B focuses specifically on the components that interact with LLMs.
According to Shostack, CISOs face a tension between the pressure to quickly secure AI systems and the difficulty of applying existing security tools to those systems. He advocates deliberately keeping threat modeling efforts low, for example through formats that fit into a one-hour meeting or a ten-minute conversation with an executive. Jeff Williams, founder of OWASP and CTO of Contrast Security, puts the urgency into perspective: AI hasn’t broken threat modeling, but rather exposed weaknesses that already existed. The practice has traditionally relied on surveys, questionnaires, interviews, outdated Visio diagrams, and spreadsheets – models that are already incomplete before a project is finished and become obsolete with every application change.
The key difference from classical software lies in the non-determinism of generative and agentic AI systems. Whereas conventional software follows rules that can be traced from input to output, AI systems interpret natural-language instructions and generate probabilistic responses – identical requests can produce different results. If agents are additionally granted access to data or tools, these results can trigger actions elsewhere. Brian Glas, Vice President of Consulting Services at CODIFIC and project lead for the OWASP Top 10, emphasizes that traditional threat modeling was designed for generally deterministic systems and that risk assessment must be adapted accordingly for modern generative and agentic AI systems.
For CISOs, the practical benefit of approaches like PHANTOM-B lies in creating a low-threshold entry point for assessing the risk of AI components, without depending on complete, time-intensive threat modeling processes. Given the increasing spread of AI applications within organizations – often outside established development and approval processes, as in the vibe-coded tool case described – a fast, repeatable initial assessment is becoming increasingly important.
Source: www.csoonline.com · Published August 19, 2026
Lumi AI News — AI-assisted curation in accordance with Art. 50 EU AI Act. Paraphrasing and classification by Lumi News Pipeline v1.8.3.