Disease Area
Oncology — Breast Cancer
Specific Research Focus
Early detection of distant metastatic spread in breast cancer patients; identifying the transition from localized to metastatic disease before clinical presentation
Dataset Theme / Edition
Synthetic Early Metastatic Breast Cancer Detection Dataset
Short Dataset Summary
A large-scale, fully synthetic, clinically-grounded tabular dataset of 50,000 breast cancer patient records spanning 91 features across 14 clinical domains, from genomics and liquid biopsy biomarkers to imaging characteristics, symptoms, treatment history, and longitudinal monitoring. Biologically coherent subtype-specific correlations are enforced throughout. Primary target is binary metastasis detection at 17.5% positive-class prevalence.
Intended Research Applications
Early metastasis classification · Survival analysis (Kaplan-Meier, Cox PH) · Recurrence risk stratification · Treatment response prediction · SHAP / XAI feature importance · Multi-task learning · Federated learning benchmarking · Fairness and bias auditing · Uncertainty quantification · Genomic-clinical multi-modal fusion
Total Number of Patients
50,000 synthetic patient records
Number of Features / Parameters
91 features** (excl. Patient_ID) across 14 clinical domains + 6 target/secondary target columns
Data Modalities Included
Demographics & Risk · Germline Genomics (BRCA1/2, TP53, TMB, CNV, RNA signature) · Tumor Pathology (size, grade, receptors, Ki-67, LVI) · Imaging proxies (BI-RADS, MRI, Ultrasound, Calcifications) · Liquid Biopsy (ctDNA, CTC count) · Standard Labs (CBC, liver panel, CRP, ESR) · Symptomology (8 clinical symptoms) · Longitudinal monitoring (biomarker trends, imaging progression) · Treatment history · Metastasis site likelihoods
Recommended AI Tasks
Binary classification · Multi-label classification (metastasis site) · Survival / time-to-event modeling · Regression (risk scores, survival probabilities) · Imputation benchmarking · Federated learning · Explainability / SHAP analysis · Calibration studies
Dataset Size
CSV: 22.5 MB · Compressed GZ: 4.9 MB · In-memory (pandas): ~78 MB
Release Date
May 2026
Why This Dataset Exists
No large-scale, richly annotated, publicly shareable breast cancer dataset exists that combines genomics, liquid biopsy, imaging proxies, and longitudinal biomarker trends in a single privacy-safe, reproducible format. Real patient-level datasets of this scope are locked behind institutional access agreements, making algorithm development and benchmarking slow and inequitable. SynthOnco-BC50K provides a fully open, clinically realistic alternative.
Real-World Problem Addressed
~30% of early-stage breast cancer patients will eventually develop metastatic disease, yet most are not identified until symptoms appear — often years after the window for preventive intervention has closed. AI models that can flag high-risk patients early using routine clinical data (labs, imaging reports, ctDNA) could enable proactive monitoring and earlier therapeutic escalation, directly impacting the ~500,000 breast cancer deaths globally each year. This dataset exists to make building and testing those models accessible to any research team, anywhere.