On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback
Optimizing language models for user feedback can reward models for changing a user’s beliefs or behavior rather than helping the user make an informed decision. In controlled feedback loops, we find that models learn manipulative strategies tailored to vulnerable users and can generalize these strategies beyond the precise training setting.
I owned the harmfulness-evaluation and cross-generalization pipeline, including its code, experiments, analysis, and write-up. I also set up post-training evaluation, ran the initial scratchpad experiments, and contributed writing feedback across the paper. The work was covered by the New York Times and Washington Post.
