Abstract: Right to Counsel policies, which provide free lawyers to tenants facing eviction, have expanded rapidly despite limited and conflicting evidence on whether and how lawyers help tenants. In a randomized controlled trial in Memphis (N = 1,140), lawyers reduce court eviction judgments by 25 percentage points (37%) on average. However, these impacts fall by 86% and become indistinguishable from zero when a concurrent rental assistance program ends. The decisiveness of rental assistance for lawyers’ impacts points to a broader lesson: a meta-analysis of prior work shows that lawyers’ integration with out-of-court rental assistance programs largely reconciles earlier studies’ mixed findings. Yet, lawyers create value for tenants even when their courtroom effects are limited. Incentivized surveys find that tenants value lawyers at twice the cost of provision — and especially for what lawyers do outside court, rather than their in-court impacts. Together, these findings that eviction lawyers’ value lies primarily out of court — both in their interaction with rental assistance programs and what tenants actually demand — clarify the rationales that would justify Right to Counsel programs and underscore how access-to-justice policies reach beyond the courtroom. [Survey Insturments]
"Racial Disparities in Forensic Labs" (draft under review, available upon request)
Abstract: Forensic laboratories sit at a critical but understudied stage of the criminal legal system. They convert physical evidence into the information that law enforcement, prosecutors, and courts rely on to move cases forward—and they do so under chronic backlog pressure. That pressure creates the need for triage, which is particularly concerning given the substantial literature documenting racial disparities in the administration of policing, charging, bail, and sentencing. Yet we know remarkably little about what happens inside forensic labs, and whether the processing of forensic evidence introduces racial disparities of its own.
This Article provides the first large-scale evidence on that question. Using case-level administrative records from the Virginia Department of Forensic Science covering more than 300,000 examination requests from 2021 to 2024, I measure processing-time disparities by the race of suspects and victims. The analysis moves from aggregate to within-jurisdiction and within-examination-type comparisons—a progression that exposes how compositional differences can mask or reverse aggregate patterns.
Two results emerge. For suspect race, cases with Black suspects appear to be processed faster in raw comparisons, but that difference disappears when comparing cases within the same jurisdiction and examination type. The aggregate pattern reflects a racially correlated distribution of forensic capacity—Black suspects are disproportionately processed in faster jurisdictions—rather than differential treatment of comparable cases. Cases with Black victims show a different pattern. Raw comparisons suggest dramatically faster processing, but within like-for-like comparisons the gap reverses: Black-victim cases are processed more slowly by roughly two days. This is Simpson's Paradox operating in a governance setting that threatens oversight: aggregate statistics suggest advantage where, in fact, comparable cases are disadvantaged. The victim disparity is distributed broadly across examiners and has grown from near zero in 2021 to approximately 7 days in 2024.
These findings illuminate how a resource-constrained evidentiary institution can generate racially patterned outcomes at two levels: through the distribution of capacity across jurisdictions and through the treatment of cases within comparable workstreams. The Article develops a measurement framework that distinguishes these channels and draws out implications for equal protection, the legitimacy of forensic triage, and the governance of administrative institutions. The performance and equity of forensic triage systems under scarcity cannot be assessed through the aggregate turnaround metrics on which oversight currently relies.
Abstract: Tax authorities use audits to detect and deter tax evasion. In practice, they commonly rely on quantitative predictions about taxpayers’ noncompliance to inform their decisions about which taxpayers to audit. We study the problem of optimal audit selection in this context. Specifically, we investigate how variation in the distribution of predicted noncompliance across taxpayers, including the uncertainty in quantitative predictions, shapes optimal audit policy. Our results highlight the contribution of these factors through a sufficient statistics characterization of the optimal audit selection rule. We leverage this characterization to quantify the social welfare benefit of varying the information available to the tax authority, for example by expanding third-party information reporting.
Abstract: Political debates often invoke “rights” to justify public transfers (e.g., the right to health care), whereas economists use welfarist frameworks which evaluate transfers’ impacts based on how they affect people’s utility. We conduct real-stakes online experiments that isolate non-welfarist from welfarist motives, and find sizable non-welfarist preferences to provide health care and legal aid to the indigent. 73% of participants make choices which are incompatible with welfarism. Non-welfarist concerns are weaker but still pervasive with neutral comparison goods. Additional experiments highlight drivers of non-welfarist motives and a key policy implication: non-welfarist concerns make Social Welfare Functions less progressive. [Survey Insturments]
Abstract: Many institutions depend on reasoned discourse to reach decisions, but the degree to which debates are publicly observable varies. We examine reasoned discourse in the U.S. Senate, and study how increasing transparency through the introduction of C-SPAN changed legislative discourse. We find that the introduction of C-SPAN encouraged members to herd with co-partisans and to anti-herd with cross-partisans; it also appears to have led to the restructuring of Senate time to facilitate performative speech. Suggesting the information problems and career incentives at play, these effects are strongest for those closest to an election and for those with less sophisticated constituencies.
"AI Assistance for Court Review of Default Judgments" with Theodora Worledge, Othman Koraichi, Daniel Bernal, Tatsunori Hashimoto, Carlos Guestrin, and David Freeman Engstrom. Accepted at 9th AAAI/ACM Conference on AI, Ethics, and Society (AIES).
Abstract: Overwhelmed courts in the United States review millions of default judgments each year. Unfortunately, such manual reviews are time-consuming and prone to error. In an audit of 188 debt collection cases granted default judgment by the Superior Court of Los Angeles, we find that 4% contained major defects that should have entirely prevented default judgment, 10% contained inconsistencies requiring reduced judgments, and 32% contained errors requiring amendment prior to judgment. To support courthouses in default judgment review, we collaborated with courthouse attorneys and judges in designing a Default Assistant. The Default Assistant employs large language models to evaluate a case with respect to predetermined legal requirements and provide cited recommendations for an expert user’s review. We equip users to verify these recommendations by grounding the assistant’s explanations in cited quotes and tables from the original case filings. We conduct a controlled study with 66 law students that conservatively simulates court review, with more time and resources than court staff. We nevertheless find users aided by the Default Assistant were 6.0% more accurate on the average requirement than unaided reviewers (p < 1.0e-4). Simultaneously, users were 25.9% faster in reviewing the average requirement than unaided reviewers (p < 2.5e-10). Statutory requirements demanding extensive document search realized the largest gains, with error reductions and time savings from AI assistance up to 62% and 34%, respectively, relative to unassisted user performance and with differences statistically significant (p < 0.05). Our work provides a proof-of-concept that AI assistants with citations have the potential to help resource-constrained courts conduct default judgment review more accurately and efficiently
"Overworking Public Defenders" conditionally accepted at American Economic Journal: Economic Policy.
Abstract: Most U.S. criminal defendants are represented by public defenders (PDs), who consistently face higher caseloads than recommended by professional guidelines. I study the effect of caseloads using novel data from three U.S. counties and quasi-random variation in case assignment timing. Higher caseloads do not change conviction rates but lengthen sentences significantly: shifting a PD from the 25th to 75th percentile of their caseload increases an average sentence by roughly 70%. PDs facing high caseloads maintain time spent on high-severity felonies at the expense of lower severity cases. These results suggest counties may realize substantial cost-savings on incarcerations by hiring additional PDs.
"Measuring discourse by algorithm" with Edward Stiglitz. International Review of Law and Economics, 2020. [PDF] [Journal Webpage]
Abstract: Scholars increasingly use machine learning techniques such as Latent Dirichlet Allocation (LDA) to reduce the dimensionality of textual data and to study discourse in collective bodies. However, measures of discourse based on algorithmic results typically have no intuitive meaning or obvious relationship to humanly observed discourse. Such measures of discourse must be carefully validated before relied on and interpreted. We examine several common measures of discourse based on algorithmic results, and propose a number of ways to validate them in the setting of Federal Open Market Committee meetings. We also suggest that validation techniques may be used as a principled approach to model selection and parameterization.
Abstract: Videoconferencing has recently become ubiquitous due to the COVID-19 pandemic but has been growing in importance for decades. Despite this growth, we have limited understanding of the costs associated with adopting this technology. In this paper I leverage a novel dataset tracking 1.7 million individuals attending 1.2 million videoconference meetings over 6 months to evaluate individual punctuality in the remote workplace. I find that participants spend a significant amount of time waiting for others to arrive. An average meeting causes 6 minutes of per participant waiting time and even small meetings (≤ 5 participants) waste 14 minutes of total participant time. I investigate the predictors of these coordination failures and find that punctuality is best (and waiting time is minimized) for smaller, shorter meetings scheduled on the hour and half hour. I find some evidence for the development of norms that lead to relatively lower coordination failures than the distributions of arrival times might suggest and discuss the implications of these findings to a time of many new users of this technology.