Skip to content

SLPO: Outcome-Reward Training for Latent Reasoners Without Token Decoding

The gist: Surrogate Latent Policy Optimization enables efficient outcome-reward training for latent reasoners that use continuous vectors instead of tokens for intermediate steps.

Researchers have developed SLPO (Surrogate Latent Policy Optimization) to transfer reinforcement learning with verifiable rewards to latent reasoning models. The method bypasses the high computational costs of explicit chain-of-thought approaches by processing intermediate steps as continuous vectors rather than fully decoding them.

Latent reasoners encode intermediate steps as continuous vectors instead of tokens, allowing them to compete with or even outperform explicit chain-of-thought procedures at shorter reasoning paths. The problem: while explicit CoT models could be scaled through outcome-reward RL, latent reasoners remained limited to imitation learning because their trajectories lack tractable step-by-step likelihoods and do not offer an adaptive stopping interface under fixed computational budgets.

SLPO addresses this limitation through two core components: an empirical surrogate policy density over latent transitions for trajectory-level reward assignment, and a stopping head with correctness supervision, which outcome-reward training refines into a variable-horizon policy. This makes reward RL applicable to autoregressive latent reasoners for the first time.

In experiments with continuous and soft-thinking settings, SLPO improves pass@k rates under parallel sampling and allocates longer latent reasoning paths adaptively for harder instances, while simpler problems are solved more accurately with shorter reasoning steps.


Source: arxiv.org · Published 21 July 2026
Lumi AI News — AI-assisted curation pursuant to Art. 50 EU AI Act. Paraphrase and classification through Lumi News Pipeline v1.7.3.

Share on: