Quantitative Research
Statistical Modeling · Machine Learning · NLP · Sentiment Analysis· Financial Analysis · Portfolio Optimization
Statistical Modeling · Machine Learning · NLP · Sentiment Analysis· Financial Analysis · Portfolio Optimization
Constructed large-scale datasets with 22,000+ observations for exploratory and quantitative analysis, producing statistical and geographic visualizations
Sourced, merged, and cleaned public datasets using Python (Pandas, NumPy), with additional data collected through API extraction and web scraping (BS4)
Created new metrics and conducted data visualization, descriptive statistics, and OLS/multiple regression analysis
Applied Lasso, Ridge, Elastic Net, clustering, Random Forest, and XGBoost with hyperparameter tuning
Solo-authored the research paper in Overleaf with LaTeX and used GitHub for version control
Skills: Pandas, Numpy, BS4, OLS, Lasso, Ridge Regressions, Clustering, Random Forests, XGBoosting, API Calls, Web Scraping, GitHub Version Control
Collaborated in a 4-person team to analyze 2023–2025 Arrive Ready email engagement, examining open and click rates, content, timing, academic streams, and link interactions
Applied descriptive statistics, OLS/ridge regression, cross-validation, and clustering in Python to identify key drivers of engagement
Combined quantitative analysis, student feedback, secondary research, and benchmarking to develop recommendations and redesign the email template
Skills: Python, Pandas, Numpy, OLS, Multiple Regressions, Ridge Regressions, Clustering (with Silhouette Score and Elbow Method), Qualitative Analysis, Secondary Research
Analyzed student survey data from Arrive Ready webinars using Python, examining satisfaction, engagement, organization, perceived length, and open-ended feedback
Applied descriptive statistics, data visualization, text cleaning, word-frequency analysis, and sentiment analysis to identify patterns in student feedback and webinar engagement
Collaborated with the Arrive Ready team to translate quantitative and qualitative findings into insights for improving future webinar content and delivery
Skills: Python, Pandas, NumPy, Matplotlib, Seaborn, TextBlob, Natural Language Toolkit (NLTK), Sentiment Analysis, Text Analysis, Word Frequency Analysis
Analyzed 12,700+ Lending Club loan records to identify key borrower and loan characteristics associated with default risk
Cleaned and transformed data by addressing missing values, encoding categorical variables, and standardizing features for machine learning models
Built and compared seven classification models, including Logistic Regression, Random Forest, AdaBoost, and XGBoost, using cross-validation and GridSearchCV for model optimization
Evaluated models using accuracy, precision, recall, confusion matrices, ROC/AUC, and training time; XGBoost achieved the strongest performance with 77.71% test accuracy
Analyzed feature importance and optimized classification thresholds based on the financial costs of false predictions, recommending a 0.3 threshold for XGBoost
Skills: Python, Pandas, NumPy, Scikit-learn, XGBoost, Cross-Validation, GridSearchCV, Data Visualization, ROC/AUC Analysis
Analyzed financial statement data for 170+ public companies, evaluating profitability, efficiency, leverage, and return on equity across industries using the DuPont framework
Built interactive dashboards to compare industry and company performance, identifying trends in financial ratios and changes over time
Skills: Excel (Power BI, PivotTables, XLookup), Tableau, financial statement analysis, DuPont analysis
Analyzed monthly financial data (Sep, 1981 - Oct, 2024) for two equities, evaluating return, risk, and performance using descriptive statistics and risk-adjusted metrics
Built and compared four portfolio optimization models (Statistical, Shrinkage, Wiener, and Itô) to determine optimal asset allocations
Skills: Excel (XLOOKUP, VLOOKUP, Data Analysis ToolPak), portfolio optimization, financial analysis, descriptive statistics, risk analysis