Title: Incentivizing Quality AI with Statistical and Adaptive Contracts
Abstract: Rethinking LLM pricing calls for an integration of ideas from machine learning, contract design, and statistics. The moral hazard introduced by current pay-per-token pricing in LLM deployments might incentivize agents to strategically route requests to cheaper, lower-quality models behind the scenes in order to cut internal inference costs. To resolve this misalignment, we introduce a pay-for-performance pricing framework grounded in contract theory, utilizing a new notion of cost-robust contracts. We establish a direct connection between the theory of optimal contract design and the theory of optimal hypothesis testing in statistics, showing that optimal contracts are essentially scaled optimal hypothesis tests, and proving a direct correspondence between statistical and economic objectives. Pay-for-performance pricing requires performance evaluation, which in our context is quality evaluation of the LLM-generated response. Such evaluation is not easy and is certainly not without cost. The notion of adaptive contracts allows detailed evaluation to be performed selectively. We provide a theoretical analysis of the tradeoff between coarse and refined evaluation, and an empirical demonstration of the benefits of adaptivity using question-answering and code-generation datasets.
** Based on joint works (in NeurIPS'24 and ICML'26) with Eden Saig, Ohad Einav, Jamie Tucker Foltz, Tamar Garbuz, and Ariel Procaccia.
Bio: Inbal Talgam-Cohen is an Associate Professor of Computer Science at Tel Aviv University. Her research lies at the intersection of theoretical computer science, economics, and law, with a focus on algorithmic game theory, incentive design, and algorithmic contracts. She received her PhD from Stanford University and also holds an LLB in Law. Her distinctions include an ERC grant, a Google Research Scholar Award, and the ACM SIGecom Best Doctoral Dissertation Award.
Title: Learning in Strategic Queuing Systems
Abstract: Over the last two decades we have developed good understanding how to quantify the impact of strategic user behavior on outcomes in many games (including traffic routing and online auctions) and showed that the resulting bounds extend to repeated games assuming players use a form of learning (no-regret learning) to adapt to the environment. However, these results assume that there are no carry-over effects between rounds: outcomes in one round have no effect on the “state” of the game in the future. Many repeated games have an evolving state, resulting in direct carry-over effect, such as repeated auctions with budgets, as well as queuing systems. In this talk we will study this phenomenon in the context of a game modeling queuing systems: routers compete for servers, where packets that do not get served need to be resent, resulting in a system where the number of packets at each round depends on the success of the routers in the previous rounds. We study the required excess server capacity needed to guarantee that all packets get served in two different queuing systems (with or without buffers) despite the selfish learning behavior of the participants.
Bio: Éva Tardos is the Jacob Gould Schurman Professor of Computer Science at Cornell University. Her research focuses on algorithms and algorithmic game theory, particularly games on networks, simple auctions, and the efficiency of systems involving self-interested users. Her distinctions include the Gödel Prize, the IEEE John von Neumann Medal, and membership in the US National Academies of Sciences and Engineering.
Title: New Results on Real-Time Reasoning in Imperfect-Information Games
Abstract: There has been tremendous progress in 2-player 0-sum games over the last 23 years. The greatest scalability improvements have come from real-time reasoning (aka. subgame-solving) techniques. Subgame solving is drastically more complicated in imperfect-information games because strategies outside a subgame affect what the strategies in the subgame should be. In the first part of this talk, I will present Obscuro, the first superhuman AI for Fog-of-War chess [Zhang & Sandholm, ICLR-26], a recognized challenge problem after superhuman level was reached in no-limit Texas hold’em. Most prior subgame-solving techniques require the construction of the “common knowledge set”, making them unusable with this much imperfect information. Our new techniques do not require that. Experiments against the prior state-of-the-art AI and human players - including the world’s best - show that Obscuro is significantly stronger. In the second part of the talk, I will present two very recent subgame-solving discoveries: 1) a new equilibrium refinement for subgame solving that maintains the safety guarantee of prior techniques while performing significantly better in practice [Kubíček, Lisý & Sandholm, IJCAI-26], and 2) how to fix modern policy-gradient algorithms, which converge to equilibrium when applied to a full game but can produce highly exploitable strategies when further trained locally at run time, by extending safe subgame solving based on gadget games to RL [Kubíček, Lisý & Sandholm, MALGAI-26].
Bio: Tuomas Sandholm is Angel Jordan University Professor of Computer Science at CMU and a serial entrepreneur. His research focuses on the convergence of AI, economics, and operations research. In parallel with his academic career, he was Founder, Chairman, first CEO, and CTO/Chief Scientist of CombineNet, Inc. from 1997 until its acquisition in 2010; the company developed the world’s largest-scale generalized combinatorial multi-attribute auctions, with over $60 billion in total spend and over $6 billion in generated savings. He has also developed the leading algorithms for several general classes of game. His team is multi-time world champion in AI-v-AI heads-up no-limit Texas hold’em. Their AI Libratus became the first and only AI to beat top humans. Then their AI Pluribus became the first and only AI to beat top humans at the multi-player game. That is the first superhuman milestone in any game beyond two-player zero-sum games. He is Founder and CEO of Strategy Robot, Inc., which focuses on defense applications, and of Strategic Machine, Inc., which provides solutions for strategic reasoning under imperfect information in applications ranging from gaming to business. Since 2010, his algorithms have been running the national kidney exchange for UNOS. He also co-invented never-ending altruist-donor-initiated chains and his algorithms created the first such chain. Such chains have become the main modality of kidney exchange worldwide and have led to around 10,000 life-saving transplants. He invented liver lobe and multi-organ exchanges, and the first liver-kidney swap took place in 2019. He is Founder and CEO of Optimized Markets, Inc., which is bringing an expressive optimization-powered paradigm to ad campaign sales, pricing, and scheduling. Among his honors are the Vannevar Bush Faculty Fellowship, AAAI Award for AI for the Benefit of Humanity, Alfred Kordelin Prize, IJCAI John McCarthy Award, Engelmore Award, Minsky Medal, Computers and Thought Award, inaugural ACM Autonomous Agents Research Award, CMU’s Allen Newell Award for Research Excellence, Sloan Fellowship, NSF Career Award, Carnegie Science Center Award for Excellence, Edelman Laureateship, and the Goldman Sachs 100 Most Intriguing Entrepreneurs Award. He is Fellow of the ACM, AAAI, INFORMS, AAAS, and AAIS. He holds an honorary doctorate from the University of Zurich.
Title: Learning and Steering Strategic Agents at Scale
Abstract: AI is moving from isolated decision-makers toward large populations of adaptive agents interacting in markets, platforms, networks, and emerging multi-agent ecosystems. At this scale, classical equilibrium computation can become computationally prohibitive, due to curses of multi-agency and independent learning. This talk develops a mean-field perspective to learning and steering in large-scale multi-agent systems, addressing two central questions: How can strategic agents and population learn at scale, and how can their collective behavior be steered toward desirable outcomes? We discuss when finite, heterogeneous multi-agent systems can be faithfully approximated by population-level models; when equilibrium learning is computationally and statistically tractable; and how decentralized agents can learn from local feedback. We then turn from learning to steering, examining how incentives can shape equilibrium outcomes or directly steer adaptive learning dynamics.
Bio: Niao He is an Associate Professor in the Department of Computer Science at ETH Zurich, where she leads the Optimization and Decision Intelligence (ODI) Group. She is the co-director of ETH-Max Planck Center of Learning Systems, and a core faculty member of ETH AI Center. Her research interests lie in large-scale optimization and reinforcement learning, with a primary focus on theoretical and algorithmic foundations for principled, scalable, and trustworthy decision intelligence. She is a recipient of AISTATS Best Paper Award, NSF CRII Award, SNSF Starting Grant, etc. She serves as the associate editor for SIAM Journal on Optimization, Mathematics of Operations Research, and regularly as an (senior) area chair for NeurIPS, ICLR, ICML and other machine learning conferences.
Title: Generative Adversarial Harnesses
Abstract: Meta-harness optimization treats the executable setup around an LLM (prompts, tools, retrieval, memory, state updates, and decision procedures) as an object of optimization, rather than a manually designed component. It is effective when task success can be measured directly, such as through unit tests, exact answers, or executable checks. Many open-ended LLM tasks, however, lack such metrics: essay quality, research synthesis, and long-form argumentation are difficult to score directly. We propose Generative Adversarial Harnesses, a framework for optimizing LLM harnesses in settings where direct evaluation is unavailable but high-quality reference outputs exist. Our framework casts harness optimization as a game between executable LLM agents: a discriminator harness that learns to distinguish expert outputs from generator outputs, and a generator harness that is optimized to produce outputs matching the reference distribution. The contribution is to lift adversarial optimization from model weights or prompts to full executable LLM harnesses, making meta-harness optimization applicable to open-ended domains where correctness is hard to measure but strong examples are available.
**This work is joint with Beomhan Baek, Junsoo Oh, Seonho Lee, and Junhyuck Kim.
Bio: Kangwook Lee is the Chief AI Officer at KRAFTON and Chief Technology Officer at Ludo Robotics. His research focuses on virtual and real AI agents, combining theoretical and empirical approaches to understand and improve how they learn, reason, and operate. Previously, he was a tenured Associate Professor at the University of Wisconsin–Madison and held research positions at KAIST. He received his PhD in Electrical Engineering and Computer Sciences from UC Berkeley and is a recipient of the NSF CAREER Award, the IEEE Joint Communications Society/Information Theory Society Paper Award, and an Amazon Research Award.
Title: Last-Iterate Convergence of Optimistic Multiplicative Weight Update
Abstract: Optimistic Gradient Descent Ascent (OGDA) and Optimistic Multiplicative-Weights Update (OMWU) are two very popular algorithms to solve convex/concave saddle-point problems, where OMWU is the non-Euclidean, entropic version of OGDA. It is known since the '80s that the last iterate of OGDA asymptotically converges to a saddle point in smooth problems. On the other hand, it is unknown if OMWU has the same property, and the literature presents only wrong proofs on this problem. In this talk, I show that OMWU converges asymptotically for smooth convex-concave saddle-point problems, with a small enough constant learning rate. This result does not require uniqueness, strict complementarity, an error bound, or initialization near a solution. The main new ingredient of the proof was discovered with assistance from ChatGPT, and I will describe the LLM-assisted proof process as well.
Bio: Francesco Orabona is an Associate Professor in the Computer, Electrical and Mathematical Sciences and Engineering Division at King Abdullah University of Science and Technology (KAUST), where he leads the OPTIMAL Lab. His research focuses on parameter-free machine learning, particularly online learning, stochastic optimization, and statistical learning theory. Before joining KAUST, he held positions at Boston University, Stony Brook University, Yahoo Research, the Toyota Technological Institute at Chicago, and several European research institutions.
** The talk of Prof. Lillian J. Ratliff is canceled.