Social Simulation with LLMs
Fidelity in Applications
@ COLM 2026
Fidelity in Applications
@ COLM 2026
In an era where digital interactions increasingly shape our social fabric, LLM-based social simulation offers a powerful lens to understand complex societal dynamics. Recent work demonstrates the potential of Large Language Models (LLMs) to model strategic decision-making [1, 5], emulate human-like interactions and memory [3, 5], and track large-scale social and cultural evolution [4, 7]. Scalable agent societies and simulation frameworks further showcase how electoral processes, collective coordination, and diverse worldviews can be represented within multi-agent systems [1, 5, 8], opening new opportunities to study collective behavior at unprecedented scale.
As the scope and ambition of these simulations expand, however, important methodological and conceptual challenges become increasingly salient. Beyond compelling demonstrations, questions remain around evaluation, robustness, interpretability, and empirical grounding. How can we distinguish substantive social dynamics from artifacts of prompting or model bias [2]? How do we meaningfully relate simulated outcomes to real-world social data? And how do we avoid misrepresenting or flattening identity dynamics and population heterogeneity in LLM-driven simulations [6]?
Now in its second iteration, our workshop invites researchers, practitioners, and thought leaders to explore rigorous and responsible approaches to LLM-based social simulation. Serving as a meeting point for communities spanning machine learning, social science, psychology, and policy, the workshop emphasizes validation and methodological soundness while fostering interdisciplinary dialogue on the ethical and societal implications of large-scale simulated societies.
[1] Altera, A., Ahn, A., Becker, N., Carroll, S., Christie, N., Cortes, M., Demirci, A., Du, M., Li, F., Luo, S., Wang, P. Y., Willows, M., Yang, F., and Yang, G. R. (2024). Project sid: Many-agent simulations toward ai civilization. arXiv preprint.
[2] Ashery, A. F., Aiello, L. M., and Baronchelli, A. (2025). Emergent social conventions and collective bias in llm populations. Science Advances, 11(20).
[3] Barrie, C. and T¨ornberg, P. (2025). Emergent llm behaviors are observationally equivalent to data leakage.
[4] Hu, T., Kyrychenko, Y., Rathje, S., et al. (2025). Generative language models exhibit social identity biases. Nature Computational Science, 5:65–75.
[5] Liu, Y., Liu, W., Gu, X., Rui, Y., He, X., and Zhang, Y. (2024). Lmagent: A large-scale multimodal agents society for multi-user simulation.
[6] Park, J. S., O’Brien, J. C., Cai, C. J., Morris, M. R., Liang, P., and Bernstein, M. S. (2023). Generative agents: Interactive simulacra of human behavior. arXiv preprint.
[7] Perez, J., L´eger, C., Ovando-Tellez, M., Foulon, C., Dussauld, J., Oudeyer, P.-Y., and Moulin-Frier, C. (2024). Cultural evolution in populations of large language models. arXiv preprint.
[8] Piao, J., Yan, Y., Zhang, J., Li, N., Yan, J., Lan, X., Lu, Z., Zheng, Z., Wang, J. Y., Zhou, D., Gao, C., Xu, F., Zhang, F., Rong, K., Su, J., and Li, Y. (2025). Agentsociety: Large-scale simulation of llm-driven generative agents advances understanding of human behaviors and society.
[9] Puelma Touzel, M., Sarangi, S., Welch, A., Krishnakumar, G., Zhao, D., Yang, Z., Yu, H., Kosak-Hine, E., Gibbs, T., Musulan, A., Thibault, C., Gurbuz, B. T., Rabbany, R., Godbout, J.-F., and Pelrine, K. (2024). A simulation system towards solving societal-scale manipulation. arXiv preprint.
[10] Takata, R., Masumori, A., and Ikegami, T. (2024). Spontaneous emergence of agent individuality through social interactions in llm-based communities.
[11] Wang, A., Morgenstern, J., and Dickerson, J. (2025). Large language models that replace human participants can harmfully misportray and flatten identity groups. Nature Machine Intelligence.
[12] Xue, Z., Jin, M., Wang, B., Zhu, S., Mei, K., Tang, H., Hua, W., Du, M., and Zhang, Y. (2025). What if llms have different world views: Simulating alien civilizations with llm-based agents. arXiv preprint.
[13] Zhang, X., Lin, J., Sun, L., Qi, W., Yang, Y., Chen, Y., Lyu, H., Mou, X., Chen, S., Luo, J., Huang, X., Tang, S., and Wei, Z. (2024). Election sim: Massive population election simulation powered by large language model driven agents. arXiv preprint.
Zachary Yang
Ubisoft La Forge | McGill University | Mila
Xuhiui Zhou
Carnegie Mellon University
Yunze Xiao
Carnegie Mellon University
Lynnette Hui Xian NG
Carnegie Mellon University
Université de Montréal | Mila
McGill University | Mila
Carnegie Mellon University
A Model of Multi-turn Human Persuadability Using Probabilistic Belief Tracing
Jared Moore, Noah Goodman, Nick Haber, Max Kleiman-Weiner
Activation-Steered Personas in Task-Oriented LLM Agent Simulations
Maryam Shoaeinaeini, Bernardo Rodrigues, Adib Mosharrof, A.B. Siddique, Brent Harrison
BehaviorBench: Benchmarking Foundation Models for Behavioral Science Tasks
Jin Huang, Yutong Xie, Wanli Song, Xingjian Zhang, Walter M. Yuan, Matthew O. Jackson, Qiaozhu Mei
Beyond Cooperative Simulators: Generating Realistic User Personas for Robust Evaluation of LLM Agents
Harshita Chopra, Kshitish Ghate, Aylin Caliskan, Tadayoshi Kohno, Chirag Shah, Natasha Jaques
Can LLMs Faithfully Enact Superforecaster Personas?
Andrew Robert Williams, Evan Jiang, Martin Weiss, Nasim Rahaman, Hugo Larochelle
Carryover Effect in LLM Personality Measurement
Yuehan Zhang, Xiao-Ming Wu
Classic AI as Scaffolding for LLM Social Agents
Anatole Gershman
Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction
Chenguang Wang, Ming Li, Xinyue Zeng, Zhuochun Li, Hong Jiao, Tianyi Zhou, Dawei Zhou
Disentangling Models from Personas in Heterogeneous LLM Simulations
Dani Roytburg, Daphne Ippolito
Distinguishing Governance Effects from Prompting Artifacts in LLM Pricing Simulations
Theja Tulabandhula
Doing What They Say, Not What They Reason: Locating the Faithfulness Gap in LLM Agents
Yufeng Wang
ECHO: Evidence-Calibrated Human Outcomes from Synthetic Users for Interaction Design
Dylan Xinming Hou
Everyone Conforms, No One Believes: Pluralistic Ignorance in LLM Agent Populations
Yashwanth YS
For Questions of Ought, AI Could Use Some SAGE Advice
Smitha Milli, Ratip Emin Berker, Sonja Kraiczy, Claudia Shi, Jack Kussman, Avinandan Bose, Himaghna Bhattacharjee, Edith Elkind, Ariel D. Procaccia, Maximilian Nickel
Human-Simulation Interaction: From Prediction to Exploration in LLM Agent Simulations for Policy
Huanxing Chen
Individually Sensible, Collectively Harmful: Simulating Commons Failure in LLM Societies
Yujiao Chen
Learning User Simulators with Turing Rewards
Yingshan Susan Wang, Cedegao E. Zhang, Linlu Qiu, Zexue He, Pengyuan Li, Alex Pentland, Roger P. Levy, Yoon Kim
LifeBench: Evaluating Large Language Models as Lifelong Decision-Makers in a Simulated Human Life
Ishaan Bansal, Shivank Garg
LLM-Augmented Agent-Based Simulation of Cyber Social Agents
Lynnette Hui Xian Ng, Kathleen M. Carley
LLMs Struggle to Measure What Distinguishes Students of Different Proficiency Levels: A Study of Item Discrimination in Reading Comprehension Assessment
Han Chen, Ming Li, Chenguang Wang, Yijun Liang, Dawei Zhou, Hong Jiao, Tianyi Zhou
Measuring Mutual Confirmation Bias Amplification in Simulated User-AI Dialogue: A Multi-Agent Simulation Framework
Reiko Tanabe
MicroVerse: An Instrument for Measuring Self-Authored Identity Drift in Long-Horizon Multi-Agent Language-Model Simulations
Sky Ng, Brihi Joshi, Ishan Gupta, Shirley Huang, Zonglin Di, Yun Shen, Qianfeng Wen, Yifan Simon Liu, Ruoqi Gao, Yilan Fan, Zhiwei Zhang, Muhammad Ahmed Mohsin, Yucheng Lu, Xiaoyi Liu, Heming Liu, Qianyu Julie Zhu, Hanwen Xing, Zhengyang Shan, My Chiffon Nguyen, Guanghui Min et al.
Moral Hazard in Multi-Agent Language Models
Dane Malenfant
More Context Is Not Validation: A Baseline-First Audit of Action-Grounded LLM Listener Agents
Chinmay Rawat, Taaha Kazi, Vasu Sharma
No One Wins in Nuclear War: Social Simulations of High-stakes Military Decision-making
Glenn Matlin, Isaac Song, Anthony Wen-Ming Zang, Mark Riedl
One Crowd, Three Cascades: Information Diffusion in LLM-Agent Simulations Is Governed by Serialization, Not Society
Tarun Raheja, Nilay Pochhi
Optimizing Behavioral Experiment Design with In-Silico Simulations
Sanchaita Hazra, Pao Siangliulue, Bodhisattwa Prasad Majumder, Peter Clark, Amy X Zhang, Joseph Chee Chang
PB-GRPO: Learning Socially Adaptive LLM Agents from Persona-Driven Simulation with Preference-Batched GRPO
Jingquan Wang, Jun Yin, Xu Han, Yongsheng Mei, Jie Hao, Bin Guo
PersonaEval: Persona-Based User Simulation for Evaluating Interactive Applications
Yifan Simon Liu, Qianfeng Wen, Yilan Fan, Shirley Huang, Ruoqi Gao, Jianheng Hou, Muhammad Ahmed Mohsin, Zonglin Di, Brihi Joshi, Xincheng Tan, Yucheng Lu, Xiaoyi Liu, Heming Liu, Hanwen Xing, Guanghui Min, Zhengyang Shan, My Chiffon Nguyen, Ishan Gupta, Yunze Xiao, Hannah Collison et al.
Position: AI Is Not Ready for Strategic Conflicts
Mark Riedl, Glenn Matlin
Position: Synthetic Persona Needs Explicit Grounding and Standardized Reporting
Yucheng Lu, Xiaoyi Liu, Heming Liu, Hanwen Xing, Shirley Huang, Guanghui Min, Zhengyang Shan, My Chiffon Nguyen, Ishan Gupta, Zonglin Di, Qianfeng Wen, Yifan Simon Liu, Ruoqi Gao, Yilan Fan, Jianheng Hou, Brihi Joshi, Muhammad Ahmed Mohsin, Yunze Xiao, Keyang Xuan, Hannah Collison et al. (24 additional authors not shown)
PrivacySIM: Evaluating LLM Simulation of User Privacy Behavior
James Flemings, Murali Annavaram
PSI-Bench: Towards Clinically Grounded and Interpretable Evaluation of Depression Patient Simulators
Nguyen Khoi Hoang, Shuhaib Mehri, Tse-An Hsu, Yi-Jyun Sun, Quynh Xuan Nguyen Truong, Khoa D Doan, Dilek Hakkani-Tür
Psychometric Framework for Comparing Alignment-Conditioned Behavior in Large Language Models
Krisha Arora, Anand Kumar Madasamy, Dishanth Arya
Reason-Mediated Behavioral Models for Auditing LLM Social Simulators
Atharva Pandey, Gautam Jajoo
Role Steering of Language Models for Social Simulations
Isaac Song, Mohammed Rehan Parwani, Glenn Matlin, Emile Timothy Anand, Akhil Theerthala, Arjun Chatterjee, Anthony Wen-Ming Zang, Maria Kostylew, Yonadav G Shavit, Sebastien Krier, Mark Riedl
Shall We Play a Game? Language Models for Open-ended Wargames
Glenn Matlin, Isaac Song, Yixiong Hao, Parv Mahajan, Evan Montoya, Ryan Bard, Stuart R. Topp, Anthony Wen-Ming Zang, Mohammed Rehan Parwani, Soham Shetty, Mark Riedl
Simulating Aggregate Electoral Opinion Trajectories with LLM Personas and Media Exposure
Jaywoong Jeong, Robin Na
Social Choice Foundations for Simulation-Augmented Generation
Sonja Kraiczy, Smitha Milli, Ratip Emin Berker, Avinandan Bose, Brandon Amos, Jamelle Watson-Daniels, Maximilian Nickel, Edith Elkind, Ariel D. Procaccia
Social OMNI-EPIC: Learning-Progress-Driven Curriculum Generation for Social Interaction
Huijun Mao, Huanxing Chen, Nick Haber
SparkMe: Adaptive Semi-Structured Interviewing for Qualitative Insight Discovery
David Anugraha, Vishakh Padmakumar, Diyi Yang
Synonymix: Unified Group Personas for Generative Simulations
Huanxing Chen, Aditesh Kumar
The Effect Of Warning Message Length on Language Model Simulated Response in Emergency Situations
Madeleine McDonald Gagné, Arnav Jhala
Tracing LLM Social Reasoning to Pretraining Data: A Provenance Audit for Social Simulation
Glenn Matlin, Chandreyi Chakraborty, Saehee Eom, Mika Okamoto, Rayan Castilla, Louis Jaburi, Alvin Deng, Taywon Min, Lucia Quirke, Stella Biderman, Mark Riedl
When the Generator Grades Itself: A Self-Judge Under-Flags Unsupported Outputs at Matched Scale in LLM Social Simulation
Rob Sneiderman
Aarati Noronha
Amogh Mannekote
Andrew Williams
Anna-Carolina Haensch
Atharva Pandey
Chen Liu
David Chan
Fangwei Zhong
Gautam Jajoo
Han Chen
Huao Li
Ishaan Bansal
Ishan Gupta
Ivy He
Jiale Han
Jiayuan Liu
Jinsook Lee
Karthika Arumugam
Kirandeep Kaur
Krishna Balusu
Kriti Faujdar
Lujun LI
Myungsu Kwak
Sankalp Jajee
Santosh Patapati
Sarthak Tayal
Shankar Venkitachalam
Shashank Lakshman
Shijun Lei
Siddharth Vohra
Smitha Milli
Tarun Raheja
Tirtho Roy
Xiangning Lin
Xinyue Zeng
Yanlin Li
Yifei Zhang
Yucheng Lu
Yufeng Wang
Zhenze Mo
Zihao Zheng
Zuhair Shaik
We're happy to be sponsored by 2077AI (https://www.2077ai.com/).