Survey on Mechanism Design with Inspection (DRAFT), 2026
This survey examines the literature on mechanism design with inspection, organized around Ben-Porath, Dekel and Lipman (2014) (henceforth BDL): a principal allocates a good to one of several privately informed agents, cannot use transfers, but can inspect an agent's claim at a cost. BDL can be interpreted as an instance of Townsend's costly state verification when transfers are not permitted. I discuss this connection along with two further related literatures that consider superficially similar problems with different primitives: costly signaling and evidence and disclosure games. I then survey work that extends or varies the BDL model along nearly every dimension of its primitives: the inspection technology (perfect, partial, probabilistic, noisy, capacity-constrained), the punishment structure, the presence and structure of transfers, the timing of inspection, the number and divisibility of goods, the informativeness of agents' own private information, and the underlying objective (welfare, revenue, or collective choice).
Slides GAIMMS 2026 Summer School
Slides from a mini-course on mechanism design with inspection.
Survey of Algorithmic Collusion (DRAFT), 2026
Algorithms can generate supra-competitive prices through several distinct mechanisms, but the literature often treats these mechanisms as variants of a single phenomenon called \emph{algorithmic collusion}. I argue that the literature is better understood as comprising three traditions. The first, originating in the theory of finite automata and program equilibrium, studies how memory, observability, and computational complexity affect the set of equilibrium outcomes that algorithms can sustain. The second studies learning under misspecification and shows how firms that learn from limited data can converge to self-confirming, supra-competitive outcomes even without strategic coordination. The third studies reinforcement-learning and no-regret algorithms, asking what patterns of behavior emerge from adaptive interaction and what algorithmic properties are sufficient to guarantee competitive outcomes. I argue that these traditions identify fundamentally different mechanisms---coordination, misspecification, and exploitation---that have distinct economic and policy implications. I conclude by suggesting that the central unresolved question is not whether algorithms can generate supra-competitive prices, but how competition operates in the space of algorithms chosen by firms subject to regulatory and informational constraints.
The Complexity of Benchmark Testing, 2026, with Lance Fortnow
Benchmark testing of large language models (LLMs) is, we argue, the same problem that economists and statisticians studied decades ago under the name \emph{forecast testing}. Forecast testing has three agents: Nature, a Forecaster, and a Tester. The correspondence with LLM benchmarking is exact. The Forecaster is the LLM: both are algorithms that, given the history so far---which includes both the forecaster's own past outputs and Nature's realized outcomes---produce a probability distribution over the next token. Nature is the benchmark's distribution over questions and their correct answers. The Tester is the grading procedure that checks whether the benchmark questions were answered correctly and decides whether to certify the model. Benchmarks are used both to compare LLMs and to certify them; we focus on certification, where certifying the LLM is the Tester passing the Forecaster. The name ``forecast testing'' is historical, from the original motivation of testing weather forecasters; nothing requires the forecaster's outputs to be probabilities of rain---any token sequence works. Under this correspondence, Sandroni's theorem---for any test that passes or fails in a finite number of steps and passes every forecaster whose forecasts agree with Nature, there is a probabilistic forecaster that passes with high probability even when Nature is adversarial---says that there is a machine-learning model that will pass any such benchmark with high probability knowing only how the test works, without knowing the questions. The Fortnow--Vohra theorem---when forecasters and testers are required to be computationally efficient, there is a distribution of Nature under which passing the test ignorantly requires factoring integers---says that benchmarks built on verifiable computational hardness resist such gaming, and LLMs indeed have trouble factoring, for analogous reasons. We develop this dictionary in full, reinterpret hallucination and benchmark contamination as instances of a forecaster optimizing against the tester rather than against Nature, and derive concrete design lessons: the two known escapes from manipulability, computational hardness and counterfactual conditioning, correspond respectively to benchmarks built from freshly generated, verifiably hard instances and to benchmarks that score consistency across families of counterfactual question variants.
Which Probability Matrices are Strict Correlated Equilibria? 2026, with Anthony Rodriguez and Can Kizilkale.
Coverage Guarantees as Ambiguity Sets: Coherent Decision Making under Statistical Uncertainty, 2026, with Selman Erol.
Confidence and conformal sets are used as inputs to downstream decisions. To act on such a set, a decision maker must give it a probabilistic interpretation, and that interpretation determines an ambiguity set: the probability distributions over states she regards as possible. This paper characterizes when a proffered set admits a coherent probabilistic interpretation, a property we call \emph{rationalizability}. For confidence and conformal sets, we characterize rationalizability by a family of linear inequalities, and we identify the maximal ambiguity set consistent with the reported set, the coverage requirement, and the observed marginals. We then study decision making over a rationalizable confidence/conformal set. A common practice assigns zero probability to states outside the proffered set. This replaces the original coverage requirement with one that can become incompatible with additional probabilistic information. The rationalizable set, however, preserves the original guarantee and remains compatible with information refinements. As information accumulates and the ambiguity set shrinks toward a singleton, the associated robust decision rule converges to Bayesian expected-loss minimization. This dynamic-consistency property is our normative justification for the coherent interpretation.
Posterior Envelopes and Bayesian Feasibility, 2026, with Thanh Nguyen.
We study Bayesian updating when actions reveal feasible sets of posterior beliefs rather than exact beliefs. In market segmentation, for example, demand curves are unobserved, but optimal prices reveal information about latent demand. With exact posterior beliefs, the splitting lemma requires only that the prior equal the average posterior. With posterior envelopes, average consistency is not enough: each interval of types must contain sufficient prior mass to support gaps between posterior bounds. Applied to segmentation, these restrictions characterize aggregate demand and yield bounds on price sensitivity, welfare, and counterfactual revenue. The framework also applies to auctions, admissions, and classification design.
Randomization and the Robustness of Linear Contracts, 2025, with Ashwin Kambhampati, Bo Peng, Zhihao Gavin Tang & Juuso Toikka
We consider contract design by an uncertainty-averse principal who does not know the production technology available to an agent. For such a principal, randomizing the choice of contract provides a hedge against uncertainty. We show that it is optimal to only use linear contracts and randomize the principal's share according to a log-uniform distribution. Thus, an optimal response to uncertainty may generate rich heterogeneity in contracts offered to observationally identical agents. Our saddle-point characterization of the optimal contract implies that the gain from randomization can be arbitrarily large; optimal randomization does not require commitment; and that asking the agent to report the technology cannot improve the principal's payoff.
Signaling Design, 2025, with Matteo Camboni, Mingzi Niu & Mallesh Pai.
We revisit the classic job-market signaling model of Spence(1973) introducing profit-seeking schools as intermediaries that design the mapping from candidates' efforts to job-market signals. Each school commits to an attendance fee and a monitoring policy. We show that, in equilibrium, a monopolist school captures the entire social surplus by committing to low information signals and charging fees that extract students' surplus from being hired. In contrast, competition shifts surplus to students, with schools vying to attract high-ability students, enabling them to distinguish themselves from their lower-ability peers. However, this increased signal informativeness leads to more wasteful effort in equilibrium, contrasting with the usual argument that competition enhances social efficiency. This result may be reversed if schools face binding fee caps or students are credit-constrained.
Persuasion Made Transparent, 2024
This note highlight the fact that a number of results in information design can be obtained without fuss or muss via linear programming.
The Network Effect of Agency Conflicts, 2019 (with Yiqing Xing and Wu Zhu).
Click here for a notebookLM podcast version (an accurate and excellent sales job).
We argue that the nature of firm-level agency conflicts counters the role of network structure in the propagation of shocks. These conflicts can have an effect on system-wide behavior that is both significant and different from those predicted based on network structure alone. This implies that corporate governance can play an important role in macro fluctuations. We consider a collection of firms linked through equity cross-holdings whose managers can take investment decisions in response to an exogenous shock. Prior work concludes that more integrated networks amplify shocks. We find that if managers are subject to default costs or limited liability, this effect is reversed because their investment decisions mitigate the spread of an initial shock. In the face of moral hazard, however, their investment choices amplify an initial shock. In particular, when the network is fully diversified the aggregate effect of idiosyncratic shocks is not small as received wisdom would suggest.