Skip to content

Hebbia Leverages Claude 3.5 for High-Precision Financial Analysis with Citation Accuracy

The Point: Claude 3.5 demonstrates a 20 percent accuracy improvement in Hebbia’s finance benchmark for financial analysis and more precise source attribution, which is critical for institutional financial due diligence.

Hebbia, an AI system for institutional financial due diligence, has tested Claude 3.5 against its finance-specific benchmark and achieved the highest accuracy values in the company’s history. In question-answer analysis over financial documents, the model achieved a 20 percent relative accuracy gain compared to its predecessor.

Hebbia serves more than one-third of the top-50 asset managers as well as Tier-1 investment banks and law firms. In this environment, clients make decisions based on analyses of thousands of dense documents, where a single wrong number can change the outcome of an entire deal. The platform combines meta-prompting with Claude models to translate natural-language queries into structured analyses across hundreds of documents. Each result lands in its own cell in Hebbia’s matrix, enabling complete transparency, traceability, and controllability.

Hebbia’s Applied AI Research team, led by Adithya Ramanathan, systematically tests each new model against a proprietary, deliberately rigidly calibrated finance benchmark. The team conducts head-to-head comparisons in which new models compete against the ones they would replace. Researcher Joe Renner conducts these tests, including question-answer and citation-finding tests over financial documents as well as agent runs with tools available in the chat product. These simulate the open, multi-source analysis that customers actually perform.

Claude 3.5 achieved the highest values to date in both test categories. In question-answer and citation analysis, the model achieved approximately 20 percent relative accuracy improvement on financial documents compared to Claude 3 Opus. Citation matching remained stable, suggesting that the gain stems from improved understanding of the evidence found. Renner observed that the model became notably more transparent: it kept every part of a multi-part query in view simultaneously, independently triggered sub-agents and tools to retrieve the correct facts, and anchored every assertion in its source rather than inferring it.

The core finding lies in two fundamental capabilities: the ability to find relevant information from dense datasets and to synthesize it correctly. These model properties have massive implications for finance and research workflows, where customers make investment decisions at scale based on the analyses generated in Hebbia.


Source: claude.com · Published July 12, 2026
Lumi AI News — AI-assisted curation in accordance with Art. 50 EU AI Act. Paraphrase and classification via Lumi News Pipeline v1.7.3.

Share on: