Disease Area
Thoracic Oncology
Specific Research Focus
Lung Cancer Early Detection in Lifelong Never-Smokers
Dataset Theme / Edition
Never-Smoker Lung Cancer — Cross-Sectional + Longitudinal Multimodal Dataset
Short Dataset Summary
10,000 cross-sectional patient records + 120,000+ longitudinal visit rows for binary lung cancer classification in never-smokers (pack-year < 1). Encodes CT imaging findings, Lung-RADS risk scores, EGFR/ALK/ROS1 driver mutations, liquid biopsy biomarkers, environmental carcinogen exposures (radon, air pollution, cooking fuel), and 5-archetype temporal disease trajectories. Prevalence calibrated to 10%, consistent with never-smoker LC incidence in mixed environmental/pulmonology screening populations.
Intended Research Applications
LC risk classification · Early nodule detection AI · Lung-RADS model benchmarking · EGFR/ALK/ROS1 biomarker research · Environmental exposure analytics · Temporal sequence modelling (LSTM/Transformer) · TimeGAN synthetic data generation · SHAP attribution studies
Total Number of Patients
10,000 cross-sectional · 120,114 longitudinal visit rows
Number of Features / Parameters
51 total — 32 cross-sectional (demographics · exposures · symptoms · CT imaging · genomics · outcomes) + 19 temporal sequence features per visit
Data Modalities Included
Structured EHR-style tabular · CT imaging metadata (findings, nodule size, Lung-RADS) · Genomic/molecular markers (EGFR, ALK, ROS1) · Liquid biopsy composite score · Environmental exposure indices · Longitudinal visit sequences · Padded NumPy tensors (10,000 × 18 × 17) with attention masks
Recommended AI Tasks
Binary classification · Sequence modelling (LSTM, Transformer) · TimeGAN temporal synthetic generation · Survival analysis · Multi-label biomarker prediction · Federated oncology model training
Dataset Size
34 MB longitudinal CSV · 14 MB dynamic features CSV · 1.4 MB static features CSV · 12 MB NumPy tensor (.npy) · 3 KB metadata JSON
Release Date
May 2026
Why This Dataset Exists
No publicly available synthetic dataset exists specifically for never-smoker lung cancer with Lung-RADS-aligned CT features, EGFR/ALK/ROS1 mutation profiles with ethnic enrichment, environmental carcinogen pathways, and longitudinal temporal trajectory sequences — the combination required to train clinically meaningful early detection AI for this underserved population.
Real-World Problem Addressed
Never-smoker lung cancer accounts for 10–25% of all lung cancer cases globally and disproportionately affects women and East Asian populations through distinct molecular pathways. These patients are frequently missed by standard smoking-history-based screening criteria and lack dedicated AI tools trained on their specific clinical presentation — leading to late-stage diagnosis in a population where early detection dramatically improves survival.