Abstract
How should we align AI systems when people genuinely disagree about how those systems should behave? In this talk, I will discuss challenges that arise at several different parts of an end-to-end collective alignment process, from theory and definitions, to gathering and acting on people’s input, to applying behavioral guidance in agentic systems. I will first examine what it means for a model to be pluralistically aligned and how different forms of pluralism can be defined and evaluated. I will then discuss how Collective Alignment contributed to OpenAI's Model Spec, including how we gathered human preferences about model behavior and translated their feedback into behavioral guidance. Finally, I will discuss ongoing work on designing systems for when agents must interpret a model spec or constitution in subjective, open-ended problems.
Bio
Mitchell L. Gordon is an Assistant Professor at MIT EECS/CSAIL and a Member of Technical Staff at OpenAI. His research designs interactive systems, models, and evaluations for AI alignment and safety. His work has been recognized with best paper awards at CHI and CSCW, an oral at NeurIPS, and Stanford’s Arthur Samuel Award for Best PhD Dissertation in Computer Science. He holds a PhD in computer science from Stanford.
Talk Recording