A lightweight adapter layer reads hidden generation states from frozen LLMs, reducing requests to larger models by up to 90.7% while maintaining performance.
Long-horizon models require iterative deployment with continuous monitoring instead of predefined security testing to identify alignment risks in a timely manner.