Common reconstruction tests for AI explanations allow models to learn false codes that produce high reconstruction scores without making individual statements verifiable — RECAP training with additional auditing heads structurally solves the problem.
Linear probes for deception detection in LLMs function reliably only on training data, not under stylistic variations—but style augmentation can restore robustness.