Skip to content

NeuroCogMap: Mapping the Functional Organization of LLMs

In brief: NeuroCogMap maps the internal representations of LLMs onto functional systems, mechanistically identifies failure patterns such as hallucinations and bias, and simultaneously improves prediction of human brain activity.

Researchers have developed a neuroscience framework called NeuroCogMap that divides internal features of large language models into functional modules, thereby explaining behaviors, failure patterns, and relationships to human cognition.

NeuroCogMap follows an approach from cognitive neuroscience: the framework organizes internal features of language models into semantically coherent functional parcels and assigns these to interpretable cognitive functions. These modules form a stable structure that is partially conserved across models and is functionally linked to the model’s outputs.

The framework enables the identification of specific internal signatures for core LLM failure patterns: hallucinations, bias, refusal errors, and sycophancy each correspond to distinct disruptions in representation and behavioral control systems. This opens up mechanism-driven approaches to detection and targeted intervention in these problems—rather than merely treating their symptoms.

Beyond model behavior, NeuroCogMap demonstrates improved predictive power for human cortical responses during natural language processing, with strongest alignment in higher-order association areas of the cortex. The framework thus offers both a system-level tool for interpreting artificial systems and an empirical bridge to human neural and cognitive functioning.


Source: arxiv.org · Published 1 July 2026
Lumi AI News — AI-assisted curation in accordance with Article 50 EU AI Act. Paraphrase and classification by Lumi News Pipeline v1.7.3.

Share on: