Selected methods and data practices for empirical corporate finance and related research.
The research design should align the question, data-generating process, comparison, estimation, and interpretation.
Define. Specify the economic mechanism, outcome, unit of analysis, and estimand.
Construct. Link sources, preserve provenance, validate identifiers, and document sample formation.
Compare. Translate identification into a plain-language counterfactual and observable diagnostics.
Stress-test. Examine magnitude, robustness, heterogeneity, uncertainty, and alternative explanations.
Selected methods and recurring data problems in empirical corporate finance.
Longitudinal firm- and establishment-level panels assembled from financial, accounting, governance, labor, policy, and geographic sources.
Issues: changing identifiers, duplicates, timing conventions, mergers, sample selection, and reproducible linkages.
Public campaign-finance records for studying political connections, policy exposure, and firm behavior.
Issues: contributor matching, employer names, committee records, transaction corrections, and policy timing.
Geographic links among firms, establishments, communities, environmental measures, labor markets, and policy jurisdictions.
Issues: coordinate systems, address quality, spatial boundaries, distance measures, and boundary changes.
Market reactions and other changes around policy or institutional events, with defined event windows and comparison groups.
Focus: timing, anticipation, parallel trends, heterogeneous treatment effects, and inference.
Event-specific comparison samples and exogenous variation when conventional panel comparisons are insufficient.
Focus: sample construction, instrument strength, exclusion restrictions, estimand interpretation, and clustered uncertainty.
Trade, quote, volatility, and options data for measuring market response, liquidity, and forward-looking expectations.
Issues: timestamp alignment, market microstructure, cleaning filters, trading calendars, and horizon definitions.
Flexible prediction methods compared with transparent statistical benchmarks and out-of-sample tests.
Focus: leakage prevention, tuning, cross-validation, feature stability, and economic—not only statistical—value.
Tests of whether information improves predictions across market regimes and forecast horizons.
Focus: real-time information sets, rolling windows, nested-model comparisons, loss functions, and out-of-sample evidence.
Design before estimation. Write down the comparison and identifying assumptions before selecting a software command.
Document the sample. Make inclusion, exclusion, matching, timing, and missing-data decisions auditable.
Interpret magnitude. Connect estimates to economic scale, institutional context, and plausible mechanisms.
Test alternatives. Use diagnostics, falsification tests, sensitivity analysis, and competing explanations.
Separate fit from evidence. For prediction, reserve information for honest evaluation. For causal claims, defend the counterfactual.
Preserve reproducibility. Version code, retain metadata, record dependencies, and make the final workflow runnable from documented inputs.
Related page: Data & Research Resources contains software documentation, data-source descriptions, and methods references.