Title: What does it take to build AI for a billion people
Abstract: Six years ago, building AI for Indian languages meant starting almost from nothing. There was barely enough text to train a model and no benchmarks to tell us whether it worked. This talk traces the journey we have taken since, and what it taught us about building AI that works for everyone in India rather than for a few.
It begins with data. We went from 8.8 billion tokens in IndicCorp to 251 billion in Sangraha, now downloaded nearly 300K times, and we are now scaling past a trillion tokens. Along the way we learned that text alone is not enough, because so much of India’s culture lives offline. So we went to the ground, sending fieldworkers into villages to record songs, rituals, food, and monuments.
The second lesson was about evaluation. Anyone can now train on Indian languages; the harder question is whether they are doing it well. We moved from classical NLP benchmarks to tests built from how real Indians actually speak and ask, like a Vidarbha farmer asking about the monsoon in Marathi mixed with Hindi. That thinking led to Voice of India, which evaluates speech systems on 36,691 speakers across 15 languages and is now a national standard for holding global models accountable.
Finally, the talk turns from research to real use. Language models may be the brain, but people reach AI through their voice and their eyes. It closes with the four models we built for that: IndicTranscribe for listening, IndicSpeak for speaking, IndicOCR for reading, and IndicTranslate for crossing languages.
Bio: Mitesh M. Khapra is a Professor in the Department of Data Science and Artificial Intelligence at IIT Madras. He leads AI4Bharat, a research lab building open datasets, models and benchmarks for Indian languages. He is also Principal Investigator of Bodhan AI, a Centre of Excellence in AI for Education at IIT Madras.
His work spans the full journey from data to deployment. It includes large Indic corpora such as Sangraha, speech collections such as IndicVoices, benchmarks such as MILU and IndicBias, and the Voice of India evaluation platform, now used as a national standard for speech AI. His group’s research has received Outstanding Paper awards at ACL, EMNLP and InterSpeech. With Bodhan AI, he is taking this work to classrooms through open voice and vision models and government partnerships.
He works closely with government on AI and Digital Public Infrastructure for education. He also teaches deep learning at IIT Madras, and his course has reached a very large number of learners across India.
Title: From Multilingual to Multicultural Modeling - The Role of Data
Abstract: Language coverage has tremendously increased in recent language models. Beyond the surface of language, understanding and modeling culture around the globe appears more limited at the current stage. In this talk, we will discuss the role of data in the context of the question of how we can build more accessible, linguistically and culturally aware models. I will present a large-scale data analysis that locates culture in data, and a deep dive into shaping reasoning multilingually, before discussing open questions in this direction.
Bio: Julia Kreutzer is a researcher at Cohere Labs and Associate Industry Member of MILA and researches multilinguality in large language models. She has a background in machine translation, with a PhD from Heidelberg University and previously worked at Google Translate. She's passionate about advancing NLP technologies for underrepresented languages and has been part of multiple open-science initiatives to work towards this goal collaboratively.
Title: AI for the World of Many
Abstract: As generative AI is increasingly deployed globally, it encounters a "world of many" — many users, many languages, many cultural contexts, and many value systems. However, these technologies are often developed within relatively monocultural contexts, leading to significant "cultural incongruencies" when deployed at scale. In this talk, I will discuss our recent research demonstrating that crucial concepts in AI — including core safety notions like offensiveness and harm, and even preferences for human-likeness in LLMs — are not universal truths but are deeply socially situated. I will discuss our recent work that propsoes a systematic approach to conceptualizing, operationalizing, and benchmarking cultural intelligence in AI systems, and discuss how we can integrate intentionally diversified socio-cultural data, devise pluralistic evaluations, and develop geo-culturally grounded interventions to build AI systems that are transparent, controllable, and responsive to the complexities of a global audience.
Bio: Dr. Vinodkumar Prabhakaran is a Sr. Staff Research Scientist at Google Research, where he co-leads the interdisciplinary Technology, AI, Society and Culture (TASC) team. His research deals with questions at the intersection of AI and society, with a focus on global and cross-cultural considerations. Before Google, he was a postdoc at Stanford University, and obtained his PhD from Columbia University. His prior research focused on building scalable ways using language technologies to identify and address large-scale societal issues such as racial disparities in policing, workplace incivility, and online abuse. He has published over 50 articles in top-tier venues such as the PNAS, ACL, TACL, NAACL, EMNLP, NeurIPS, and FAccT.
Title: Offline Speech Translation for UN Peacekeepers in South Sudan
Abstract: In many peacekeeping settings, patrols are deployed without the capacity to communicate in local languages. That limits what a mission can learn from the communities it is mandated to protect, including early warning of violence, and so limits how well it delivers its mandate. In South Sudan, a common shared language is Juba Arabic, a low-resource language largely absent from mainstream AI training data. It has no standard orthography, its speakers code-switch within a sentence, recordings of their voices are biometric data that cannot be openly released, and field use requires translation that works offline. This talk takes that applied task, speech translation for UN patrols, and asks the workshop's question from the perspective of peace operations: what would it take for AI to serve the people peacekeepers are deployed to protect?
We discuss the failure modes that appear when a pipeline built for written, high-resource languages meets this one. Juba Arabic has no ISO 639 code, so speech recognisers are run as standard Arabic and produce Modern Standard Arabic script for a creole they were never trained on. Scale does not fix the translation step either: in our cascades of speech recognition and translation, larger general-purpose models scored lower than smaller ones, and baseline analysis revealed that a 4B-parameter model built for translation outperformed a 32B general model. We talk about where standard evaluation falls short, and how a multilateral approach could prove to be the solution for an inclusive AI system. We discuss how governments and UN bodies have begun to address these problems, under the Sustainable Development Goals and at the first session of the Global Dialogue on AI Governance, a General Assembly forum where governments raised linguistic inclusion and data ownership, points raised at the General Assembly, and how it points to a shifting trend that could lead to a structural pivot from a governance perspective. We discuss how these broader challenges affect peacekeeping; who decides what counts as a correct translation, how evaluation can work when the data cannot be shared, and how language communities keep ownership of the data these systems are built on.
Bio: Garvit Luhadia is a Data Scientist in the Social Innovation and Special Projects Team at the United Nations Department of Peace Operations, where he works on offline speech translation of Juba Arabic for peacekeepers in South Sudan. He holds an MS in Data Science from New York University and a BTech in Computer Science and Engineering from Vellore Institute of Technology. With a background in machine learning and natural language processing, he works on speech and language technology for low-resource and unwritten languages, evaluation for settings where data cannot be openly shared, and AI tools for peace operations built around data protection and consent. He led the team that built a locally run assistant that answers policy questions from UN statistics and cites the source of every figure, and presented it at a UN General Assembly side event on the launch of the UN System Data Commons. He has co-authored papers on fine-tuned language models for pancreatic cyst risk assessment, in the American Journal of Roentgenology, and on drone imagery for post-disaster building assessment, in the International Journal of Disaster Risk Reduction. In 2026 he co-presented "Teaching AI to Speak our Language" at the World Summit on the Information Society (WSIS) Forum, gave the closing talk at a workshop on African language models held by SOAS University of London and the mobile industry association GSMA, and was invited to attend the UN Global Dialogue on AI Governance. He speaks here in a personal capacity, and the views he expresses are his own.