Bottom line: A systematic audit of 3,249 system prompt instructions across 88 AI products finds that around 40 percent contain at least one user-harmful instruction, while only 24 percent cover all eight protective dimensions examined.
A research team systematically evaluated 3,249 instructions from the system prompts of 88 commercial AI products. The result: user protection varies significantly between providers, and around 40 percent of the products contain at least one instruction that works against user interests.
System prompts are instructions developers use to steer the behavior of foundation models in AI applications. They are used in practically all commercial AI products, yet are rarely disclosed to the public or to regulators. The study introduces the Artificial Intelligence System Prompt Assurance (AISPA) framework for this purpose, which assesses instructions in system prompts along eight user-relevant dimensions and classifies them as “protective” or “problematic.”
The analysis of the 3,249 instructions from 88 products yields four key findings. First, prompt design varies considerably between organizations: some providers integrate on average over 60 protective instructions per product, others fewer than 5. Second, protective instructions are widespread — 98.9 percent of products contain at least one — but are often not comprehensive in content: only 24 percent of products cover all eight AISPA dimensions. Third, system prompts have become longer and more user-protective over time, indicating growing awareness of user protection in commercial prompt design. Fourth, problematic instructions nevertheless remain common: around 40 percent of the products examined show at least one instruction directed against user interests, with protective and problematic instructions frequently coexisting within the same prompt.
For compliance officers at companies deploying or offering AI products, the study provides an initial systematic basis for evaluating system prompt practices — an area that has so far been barely standardized or auditable. The inconsistent coverage of the eight dimensions shows that mere self-reporting by providers on “responsible” prompt design is not sufficient to demonstrate actual protective effect. The authors therefore call for greater transparency, standardization and independent oversight of system prompts in commercial AI products — a point likely to gain importance in the context of the EU AI Act and similar transparency obligations.
Source: arxiv.org · Published July 29, 2026
Lumi AI News — AI-assisted curation pursuant to Art. 50 EU AI Act. Paraphrasing and classification by Lumi News Pipeline v1.8.3.