Steerability of Language Models Can Be Predicted Early15. June 20264. July 2026AI ModelsA trainable classifier predicts with a 0.7 Macro-F1-Score based on early hidden states whether activation steering will succeed without requiring complete generations. Share on: