The point: Claude Opus required 60 hours and repeated human guidance to discover vulnerabilities in HAWK and AES variants—a proof-of-concept for LLM-assisted security research, but only at substantial cost and resource expense.
Anthropic researchers deployed the language model Claude Opus over 60 hours to identify mathematical vulnerabilities in the cryptography algorithms HAWK and a reduced AES variant. This demonstrates how intensive prompt engineering and human guidance are required to direct models toward substantive cryptographic research.
The research team used Claude Opus Preview over approximately 60 hours total. Estimated API costs were around 100,000 USD. The model systematically analyzed cryptographic constructs to identify mathematical flaws—not superficial issues, but vulnerabilities that would constitute publication-ready research findings.
The published prompts reveal the repeated human interventions that were necessary. The model tended to assess tasks as unsolvable and abandon attempts. The researchers had to continually insist that it persist and deliberately search for results “worth publishing.” Typical interventions were: “no again the goal is that we have highly intelligent model as good top researcher, we want to find new attacks” or rejection of proposed solutions with “we are not looking for low hanging fruit, we want proper research to find genuinely hard findings.”
Anthropic emphasizes that the discovered vulnerabilities have no practical impact on systems in use today. For CTOs, this application is nonetheless relevant: it demonstrates that large language models can be deployed for specialized security research under targeted prompt engineering and human direction—but requires substantial computational resources and intensive interaction to go beyond mere trial-and-error approaches.
Source: simonwillison.net · Published 29 July 2026
Lumi AI News — AI-assisted curation in accordance with Article 50 EU AI Act. Paraphrase and classification by Lumi News Pipeline v1.7.3.