Princeton
Title: Making the Case for Small Language Models
Abstract: The talk will discuss a recent paper on LMs playing and explaining chess, as well as a retrospective on lessons from training SLMs over the years.
Speaker Details:
Danqi Chen is an Associate Professor of Computer Science at Princeton University, where she co-leads the Princeton NLP Group and serves as Associate Director of Princeton Language and Intelligence (PLI). Her research focuses on natural language processing, machine learning, and the development of intelligent systems that can acquire and reason over knowledge from large-scale text data. Her contributions have been recognized through numerous prestigious honors, including the NSF CAREER Award, and the Samsung AI Researcher of the Year.
Tatsu Hashimoto
Stanford
Title: Scaling down high-compute phenomena
Abstract: Empirical machine learning constantly faces the challenge of studying interesting high-compute phenomena (such as algorithmic aspects of high-compute pretraining or frontier capabilities) while keeping our experimental costs low enough to enable robustness and replication. While tools such as compute scaling laws have been developed for this task, the application of scaling laws to studying high-compute phenomena can be subtle and tricky. In this talk, we will discuss how careful treatment of the scaling axis and variables allows us to build scaling laws for studying high-compute, data-constrained settings as well as advanced, post-trained capabilities that are pre-emergent. In both cases, we find the scaling surprisingly predictable across models and scales, and that these approaches enable radical downscaling of our experiment designs.
Speaker Details:
Tatsunori Hashimoto is an Assistant Professor of Computer Science at Stanford University. His research focuses on developing statistical and machine learning methods that improve the robustness, safety, and reliability of large-scale AI systems. He is a recipient of the NSF CAREER Award, the Alfred P. Sloan Research Fellowship, and the Samsung AI Researcher of the Year, and has received best paper awards at ICML, ICLR, and CHI. He received his Ph.D. from Massachusetts Institute of Technology.
CMU
Title: Building predictive models for AI training
Speaker Details:
Andrew Ilyas is an Assistant Professor at Carnegie Mellon University. His research seeks to make ML more predictable & reliable by pursuing a precise empiricalunderstanding of the entire ML pipeline. Prior to joining CMU, he was a Stein Fellow at Stanford University and received his Ph.D. from Massachusetts Institute of Technology. His work has been recognized through honors including the George M. Sprowls Ph.D. Thesis Award and the Open Philanthropy AI Fellowship.
MIT
Title: AI at the last mile
Speaker Details:
Omar Khattab is an Assistant Professor of Electrical Engineering and Computer Science at Massachusetts Institute of Technology. He studies how to program and optimize systems specified in natural language, as well as how to use the knowledge encoded in language models to build more effective and sample-efficient learning algorithms. He is the creator of the ColBERT retrieval model and the DSPy framework, which have become widely used foundations for retrieval-augmented generation and language model programming.
MSR/UW-Madison
Title: The golden age of asking questions
Abstract: Over the past few months, agents have changed the way I do research by collapsing the distance between a question and an experiment. Ideas that used to die under the technical burden of setting up experiments are now testable in days, often by one person, several agents, and a laptop. In this talk I'll walk through a few side projects I ran since February: training the smallest transformer that can do 10-digit addition, predicting LLM benchmarks without running them, and building the first trained computer that is a transformer. Then I'll share what happened when I turned the AI Death Star to bigger, longstanding open questions that used to torment me as a grad student. I'll share my thoughts in what type of research agents seem to enable, and why taste and the ability to verify become more important aa execution becomes cheap and abundant. I'll close on a question I don't have a good answer to: we've always trained researchers through technical work, and taste and verification emerge from it. What happens when most of it becomes automated?
Speaker Details:
Dimitris Papailiopoulos is a Principal Researcher with the AI Frontiers lab at Microsoft Research, and the Jay and Cynthia Ihlenfeld Associate Professor of Electrical and Computer Engineering at University of Wisconsin-Madison. His research focuses on machine learning, optimization, coding theory, and the foundations of large-scale AI systems. He is a recipient of the NSF CAREER Award and the IEEE Joint Communications Society/Information Theory Society Best Paper Award, and is a co-founder of the Conference on Machine Learning and Systems (MLSys).
Thinking Machines
Speaker Details:
Songlin Yang is a Member of Technical Staff at Thinking Machines Lab. Her research focuses on efficient language model architectures, sequence modeling, and hardware-aware algorithm design, with an emphasis on linear attention methods and scalable alternatives to Transformers. She received her Ph.D. from Massachusetts Institute of Technology, where she was advised by Yoon Kim.