Artificial Intelligence

Qian Hu

Author: Qian Hu

Teaching models to forget: Selective unlearning with Amazon Nova

In this post, we introduce Reverse Direct Preference Optimization (rDPO), the novel unlearning technique behind Amazon Nova Customizable Content Moderation Settings (CCMS), and show how it reduces over-deflection while preserving model quality. We also provide pointers for customers who want to apply these preference optimization techniques to their own experiments.