In addition to PPO-based RLHF, Studio also supports DPO (Direct Preference Optimization). DPO eliminates the need for a separate reward model by directly updating the base model using preference data.
To use DPO in Studio, simply connect your Preference Dataset node directly to the DPO Optimization node.