Novel cognitive science grounded approach to preference modeling: synthetic counterfactual training + attention-based attribute integration. Impressive empirical validation across 45 communities with human evaluation confirming interpretability claims. This opens exciting research directions!
WHY do you prefer something over another? Reward models treat preference as a black-box😶🌫️but human brains🧠decompose decisions into hidden attributes We built the first system to mirror how people really make decisions in our recent COLM paper🎨PrefPalette✨ Why it matters👉🏻🧵