Disease Area:
Psychiatry / Mental Health
Specific Research Focus:
Early detection and longitudinal monitoring of Major Depressive Disorder (MDD) in adolescent populations
Dataset Theme / Edition:
Synthetic Adolescent MDD Cohort — Cross-Sectional + Longitudinal Edition, v1.0
Short Dataset Summary:
A fully synthetic, privacy-preserving dataset of 10,000 adolescent patients (ages 12–19) with DSM-5-consistent MDD diagnostic labels, 36 cross-sectional clinical and behavioral features, and 120,093 longitudinal visit records across 6–18 visits per patient. Generated with evidence-based inter-feature correlations, five clinical trajectory archetypes, and realistic missing data patterns.
Intended Research Applications:
Binary MDD classification, early school-based intervention modeling, longitudinal risk prediction, survival analysis (time-to-hospitalization), explainable AI (XAI/SHAP attribution), PHQ-9 benchmarking, federated learning across simulated clinical sites, fairness and bias auditing, TimeGAN sequence generation benchmarking
Total Number of Patients:
10,000 (with 120,093 total visit rows across the longitudinal file)
Number of Features / Parameters:
36 cross-sectional features + 20 longitudinal visit-level features = 56 total columns in the merged dataset; 57 entries documented in the data dictionary
Data Modalities Included:
Structured tabular (cross-sectional patient records), time-series sequences (longitudinal visit records), padded NumPy tensors with attention masks, static/dynamic feature splits for sequential deep learning
Recommended AI Tasks:
Binary classification (MDD diagnosis), ordinal regression (depression severity), time-series forecasting (symptom progression), survival/time-to-event modeling, anomaly detection (rapid deterioration), federated learning, generative modeling (TimeGAN-CR), fairness-constrained learning
Dataset Size:
~58 MB total across all output files — CSV longitudinal (30.4 MB), NumPy tensor (12.4 MB), dynamic features CSV (12.2 MB), cross-sectional CSV (1.5 MB), static features CSV (0.8 MB), attention mask (0.7 MB), plus metadata and data dictionary
Release Date:
May 2026
Why This Dataset Exists:
Real-world adolescent MDD datasets are extremely scarce due to strict HIPAA/GDPR protections, IRB barriers, and the sensitivity of pediatric mental health records. This synthetic dataset was created to give researchers a clinically realistic, freely usable alternative that enables model development, benchmarking, and fairness testing without requiring access to protected patient data.
Real-World Problem Addressed:
Adolescent MDD affects approximately 1 in 5 teenagers and is the leading cause of disability in young people globally, yet the majority of cases go undetected until they reach crisis severity. The absence of shareable training data blocks the development of AI-powered early detection tools in school, telehealth, and pediatric primary care settings — precisely where early intervention would have the greatest impact.