Direct-OPD: Transferring Policy Shifts from Smaller to Larger Models14. July 2026AI ModelsDirect-OPD transfers RL-induced policy shifts from weaker to stronger models by leveraging the implicit reward signal from the log-ratio of the RL-shifted and original policy. Share on: