In brief: Large language models do not consistently aggregate predictions over subpopulations into valid estimates for overall populations, despite possessing the necessary knowledge.
Researchers have examined whether large language models follow fundamental laws of probability theory in in-context learning. The finding: LLMs systematically violate the law of total probability when aggregating predictions over subpopulations to overall populations.
The research investigates whether in-context learning in LLMs can be interpreted as conditioned inference — that is, whether the model estimates a valid conditional probability distribution based on a context description. Under this assumption, LLM estimates should satisfy the law of total probability: weighted conditional distributions over subpopulations should aggregate to marginal distributions of an overall population, regardless of how the population is partitioned.
The researchers use binary trees as an evaluation framework to recursively subdivide populations into increasingly fine-grained subpopulations. They prompt LLMs with verbalized subpopulation descriptions, aggregate the estimates back to population level, and compare them across partitions of different granularity. The protocol is applied to multiple problem domains and frontier models.
The results reveal widespread violations of fundamental consistency principles. A detailed study of persona prompting reveals a phenomenon the authors term the “Macro Fallacy”: estimates reconstructed from finer-grained subpopulation responses often agree better with human reference data than direct population-level estimates. This effect persists across variations in tree structure and estimation task and can be partially offset through implicit prompting.
The findings suggest that models possess relevant subpopulation knowledge but fail to reliably propagate it in aggregate estimates. The gap establishes statistical self-consistency as a reference-free, fundamentally saturated evaluation criterion for LLMs — a test that requires no gold-standard annotations but only checks internal consistency rules.
Source: arxiv.org · Published 15 July 2026
Lumi AI News — AI-assisted curation pursuant to Art. 50 EU AI Act. Paraphrase and classification by Lumi News Pipeline v1.7.3.