Skip to content

Data Quality, Not Model Choice: What Really Determines SOC Success

In brief: A controlled study using four network telemetry sources and multiple LLMs shows that high-quality, protocol-rich security data such as Corelight logs improve core investigation metrics by two to four times – regardless of the language model used.

A controlled study finds that it is not the LLM deployed but the quality of the underlying network telemetry that decisively determines the success of AI-assisted security operations workflows. According to the results, high-quality network evidence improves key investigation metrics by a factor of two to four.

As AI increasingly takes on tasks within Security Operations Centers (SOCs) – such as threat triage, extraction of indicators of compromise, or generation of incident reports – security leaders face a fundamental question: does SOC performance depend more on the Large Language Model (LLM) deployed, or on the quality of the security data it processes? Academic studies offer conflicting answers to this question. An article in Frontiers in Artificial Intelligence emphasizes differences in accuracy, relevance, and clarity between models such as Claude 3.5 Sonnet, Gemini, ChatGPT 4o, and Mistral Large 2. A publication in the journal Information Systems, by contrast, argues that trustworthy AI applications require high-quality training and test data across multiple quality dimensions such as accuracy, completeness, and consistency.

The research project “Provably Better Data” examined this question within a controlled test framework. Two operational benchmarks were used: a capture-the-flag scenario comprising 44 questions related to a Volt Typhoon attack campaign, and an incident-response report-writing task based on a Salt Typhoon dataset. To isolate data quality as a variable, the test setup processed four different network telemetry sources under identical conditions: enriched Corelight logs, open-source nDPI firewall logs, Snort 3 IDS alerts, and NetFlow connection telemetry. Each dataset was run multiple times through a schema normalized according to OCSF (Open Cybersecurity Schema Framework). The models used were Anthropic Claude Opus 4.6, Google Gemini Pro 3.1 Preview, and older models from both providers, each with identical prompts across all test runs. Accuracy on the CTF questions and the proportion of incident-response statements substantiated by available evidence were both evaluated.

The results reveal an evidence ceiling in automated workflows: frontier models possess advanced reasoning capabilities, but the quality of their conclusions remains limited by the data available. If telemetry lacks detailed protocol-level context, AI agents cannot infer what was never captured in the first place. A concrete example from the CTF test illustrates this: when asked for the NetBIOS computer name associated with the IP address 10.110.154.113, the Corelight logs provided the correct answer, “FINANCE01” – found in the server_nb_computer_name field of the NTLM log, since Zeek had parsed the NTLM Type 2 challenge message and extracted the server’s computer name. The firewall logs did detect NTLM protocol activity but were unable to parse the individual fields within the NTLM challenge, and therefore failed to produce the correct answer.

For CISOs, this carries direct operational consequences: concrete, enriched telemetry reduces mean time to respond (MTTR), limits token consumption in LLM-driven workflows and thus their cost, and increases the demonstrable return on investment of security measures. The study results thus provide a basis for justifying investments in high-quality telemetry infrastructure to executive leadership and the board, for reducing analyst turnover caused by alert fatigue, and for presenting robust security metrics. When planning SOC architecture, the choice and enrichment of data sources should therefore be given at least as much weight as the selection of the language model deployed.


Source: www.csoonline.com · Published August 17, 2026
Lumi AI News — AI-assisted curation pursuant to Art. 50 EU AI Act. Paraphrasing and classification by Lumi News Pipeline v1.8.3.

Share on: