Multi-Head Latent Control: Reading Agent Decisions Directly from the Model27. July 2026AI ModelsA lightweight adapter layer reads hidden generation states from frozen LLMs, reducing requests to larger models by up to 90.7% while maintaining performance. Share on: