As the current LLM research community pushes towards massive scale, this workshop takes a complementary approach by focusing on the under-explored small-scale regime, where scale encompasses compute, data, and model size. We ask: to what extent is scale necessary, and how far can we push toward smaller settings while maintaining competitive performance and enabling scientifically meaningful discoveries?
Beyond transferring insights from small- to large-scale settings, we emphasize the intrinsic value of understanding small-scale regimes. Practically, studying small-scale limits can yield more efficient algorithms and system designs. Small-scale settings also enable controlled, systematic experimentation with rapid iterations, thereby facilitating scientific progress.
This workshop aims to highlight the methods, opportunities, and insights enabled by small-scale experimentation, and to foster discussion on how such approaches can broaden participation and accelerate progress in LLM research.
Date: Oct 9, 2026
Location: Hilton Union Square (Union Square 5&6 / Fourth Floor), Theater 84.
(9:00 - 9:10) Opening Remarks
(9:10 - 10:30) Invited Talks #1, #2 (2 x 40min)
Danqi Chen -- Making the Case for Small Language Models
The talk will discuss a recent paper on LMs playing and explaining chess, as well as a retrospective on lessons from training SLMs over the years.
Omar Khattab -- AI at the last mile
(10:30 - 10:45) Coffee break
(10:45 - 11:25) Invited Talk #3 (40min)
(11:25 - 12:10) Contributed Talks and Demos (3 x 15min)
(12:10 - 13:30) Lunch break
(13:30 - 14:30) Panel Discussion (-- take the poll on discussion topics!)
(14:30 - 14:40) Coffee break
(14:40 - 16:00) Invited Talks #4, #5
Tatsu Hashimoto -- Scaling down high-compute phenomena
Empirical machine learning constantly faces the challenge of studying interesting high-compute phenomena (such as algorithmic aspects of high-compute pretraining or frontier capabilities) while keeping our experimental costs low enough to enable robustness and replication. While tools such as compute scaling laws have been developed for this task, the application of scaling laws to studying high-compute phenomena can be subtle and tricky. In this talk, we will discuss how careful treatment of the scaling axis and variables allows us to build scaling laws for studying high-compute, data-constrained settings as well as advanced, post-trained capabilities that are pre-emergent. In both cases, we find the scaling surprisingly predictable across models and scales, and that these approaches enable radical downscaling of our experiment designs.
Dimitris Papailiopoulos -- The golden age of asking questions
Over the past few months, agents have changed the way I do research by collapsing the distance between a question and an experiment. Ideas that used to die under the technical burden of setting up experiments are now testable in days, often by one person, several agents, and a laptop. In this talk I'll walk through a few side projects I ran since February: training the smallest transformer that can do 10-digit addition, predicting LLM benchmarks without running them, and building the first trained computer that is a transformer. Then I'll share what happened when I turned the AI Death Star to bigger, longstanding open questions that used to torment me as a grad student. I'll share my thoughts in what type of research agents seem to enable, and why taste and the ability to verify become more important aa execution becomes cheap and abundant. I'll close on a question I don't have a good answer to: we've always trained researchers through technical work, and taste and verification emerge from it. What happens when most of it becomes automated?
(16:00 - 18:00) Poster Session & Closing Remarks
Contact: colm2026-moss-workshop [at] googlegroups [dot] com
Twitter: @MOSS_workshop