Direct-OPD transfers RL-induced policy shifts from weaker to stronger models by leveraging the implicit reward signal from the log-ratio of the RL-shifted and original policy.
Claude and other language models can handle simple robotics tasks when working through predefined controllers, but fail at direct motor control without additional abstraction.
A modified Transformer with two independent computation streams for state management and token prediction reduces required resources and improves performance by 2–3 percentage points on downstream tasks.
Sumi is the first openly available Uniform-Diffusion language model trained from scratch at the 7-billion-parameter scale and addresses a research gap between established autoregressive and masked diffusion approaches.