Unveiling Themes in 10-K Disclosures: A New Topic Modeling Perspective, 2025, International Review of Financial Analysis, 103, p.104121 (with Matthias Fengler); DOI: https://doi.org/10.1016/j.irfa.2025.104121
Sentiment-semantic Word Vectors: A New Method to Estimate Management Sentiment, 2024, Swiss Journal of Economics and Statistics, 160 (9); DOI: https://doi.org/10.1186/s41937-024-00126-1
Forecasting the Macroeconomy with Corporate Disclosures and Language Models; DOI: http://dx.doi.org/10.2139/ssrn.5594411
Abstract: This paper examines the predictive power of corporate disclosures in macroeconomic forecasting. Leveraging advanced language models, I analyze the prevalence and tonal polarities of 31 accounting topics across 890,277 10-K and 10-Q filings. Firm disclosures exhibit superior performance in long-horizon forecasts, with their predictive strength primarily driven by discussions on corporate income taxes, revenue performance, mergers and acquisitions, and repurchase agreement activities. These topics transmit information to macroeconomic outcomes through firms' capital expenditures and dividend policies. Moreover, firm disclosures provide particularly strong forecasts of GDP and Investment during periods of economic downturn, outperforming the FRED-MD dataset.
Presentation: Econometric Society European Summer Meeting (EEA-ESEM) 2026 (scheduled); International Association of Applied Econometrics Annual Conference 2026 (scheduled); 13th Nordic Econometric Meeting; 7th ECONDAT Fall Meeting 2025.
Read Them All at Once - Efficient Language Modeling for Very Long Financial Documents (with Erik-Jan Senn); Revise & Resubmit at the European Accounting Review; DOI: http://dx.doi.org/10.2139/ssrn.5011898
Abstract: This paper introduces LongFinBERT, a language model designed for excessively long financial documents. LongFinBERT demonstrates significant efficiency, requiring considerably less computational time and memory than other state-of-the-art language models. We apply LongFinBERT to two financial applications: (i) detecting financial misreporting, and (ii) examining stock market reactions to year-over-year modifications in firm disclosures. In misreporting detection, LongFinBERT outperforms accounting variables and other text-based models. In the second application, we find that investors respond to modifications in firm disclosures, as measured by LongFinBERT, with significantly stronger reactions during economic downturns. Importantly, conventional text-analysis approaches underestimate this insight.
Presentation: 47th Annual Congress of the European Accounting Association 2025; Financial Econometrics and Machine Learning Conference 2024; Financial Fraud, Misconduct and Market Manipulation Conference 2024; 17th International Conference on Computational and Financial Econometrics 2023.
Selective Disclosure and Firm Characteristics: Evidence from Cyber Breaches (with Hao Ma)
Abstract: Firm characteristics built from voluntary firm disclosures inherit a priced selection bias---disclosure risk. We document this for cyber breaches: severe losses are systematically less likely to be quantified, with CEO unearned equity incentives the dominant driver of concealment. Sorting on disclosure-risk exposure within breach-probability bins delivers a monthly FF6 alpha of 0.61%; unconditional sorts on the same measure are insignificant. The result suggests that selection bias is priced after controlling for underlying risk levels, turning firm characteristics into vehicles for the disclosure risk itself.
Presentation: Financial Fraud, Misconduct and Market Manipulation Conference 2026 (scheduled)
Trade War and Corporate FX Hedging (with Kiet Duong)
Learning Who Conceals: Double Machine Learning of Cyber-Loss Disclosure