Unveiling Themes in 10-K Disclosures: A New Topic Modeling Perspective, 2025, International Review of Financial Analysis, 103, p.104121 (with Matthias Fengler); DOI: https://doi.org/10.1016/j.irfa.2025.104121
Sentiment-semantic Word Vectors: A New Method to Estimate Management Sentiment, 2024, Swiss Journal of Economics and Statistics, 160 (9); DOI: https://doi.org/10.1186/s41937-024-00126-1
Forecasting the Macroeconomy with Corporate Disclosures and Language Models; DOI: http://dx.doi.org/10.2139/ssrn.5594411
Abstract: This paper examines the predictive power of corporate disclosures in macroeconomic forecasting. Leveraging advanced language models, I analyze the prevalence and tonal polarities of 31 accounting topics across 890,277 10-K and 10-Q filings. Firm disclosures exhibit superior performance in long-horizon forecasts, with their predictive strength primarily driven by discussions on corporate income taxes, revenue performance, mergers and acquisitions, and repurchase agreement activities. These topics transmit information to macroeconomic outcomes through firms' capital expenditures and dividend policies. Moreover, firm disclosures provide particularly strong forecasts of GDP and Investment during periods of economic downturn, outperforming the FRED-MD dataset.
Presentation: Econometric Society European Summer Meeting (EEA-ESEM) 2026 (scheduled); International Association of Applied Econometrics Annual Conference 2026 (scheduled); 13th Nordic Econometric Meeting; 7th ECONDAT Fall Meeting 2025.
Read Them All at Once - Efficient Language Modeling for Very Long Financial Documents (with Erik-Jan Senn); Revise & Resubmit at the European Accounting Review; DOI: http://dx.doi.org/10.2139/ssrn.5011898
Abstract: This paper introduces LongFinBERT, a language model designed for excessively long financial documents. LongFinBERT demonstrates significant efficiency, requiring considerably less computational time and memory than other state-of-the-art language models. We apply LongFinBERT to two financial applications: (i) detecting financial misreporting, and (ii) examining stock market reactions to year-over-year modifications in firm disclosures. In misreporting detection, LongFinBERT outperforms accounting variables and other text-based models. In the second application, we find that investors respond to modifications in firm disclosures, as measured by LongFinBERT, with significantly stronger reactions during economic downturns. Importantly, conventional text-analysis approaches underestimate this insight.
Presentation: 47th Annual Congress of the European Accounting Association 2025; Financial Econometrics and Machine Learning Conference 2024; Financial Fraud, Misconduct and Market Manipulation Conference 2024; 17th International Conference on Computational and Financial Econometrics 2023.
Cyber Risk, Strategic Disclosure, and Asset Prices (with Hao Ma)
Abstract: Cyber risk depends on both the likelihood that a breach enters the public record and the severity of its financial consequences. Yet firms retain discretion over whether those consequences are quantified. Using U.S.-listed firms from 2005 to 2025, we separate breach observability from loss quantification. We find that larger latent losses are less likely to be quantified, so visible losses understate conditional log-dollar severity. Quantification is negatively associated with CEO unearned equity incentives and salary share, consistent with strategic disclosure. We construct a firm-level disclosure-risk exposure from the reporting-margin selection correction and show that it is priced conditionally on public-record breach likelihood: within the highest breach-probability bin, portfolios sorted on reporting-margin exposure earn a value-weighted Fama-French six-factor alpha of up to 0.78% per month. The loss-quantification process therefore shapes measured cyber exposure and contains return information beyond breach likelihood.
Presentation: Financial Fraud, Misconduct and Market Manipulation Conference 2026 (scheduled)
Trade War and Corporate FX Hedging (with Kiet Duong)
Learning Who Conceals: Double Machine Learning of Cyber-Loss Disclosure