Abstract Purpose: Breast density is an important risk factor for breast cancer (BC) as well as contralateral breast cancer (CBC). However, it is missing in most publicly available datasets as well as is unknown at individual level for many women. This limits the applicability of BC and CBC risk prediction models to individual patients as well as validation of the models. The goal of this study was to develop a model to predict BI-RADS breast density using routinely collected, non-imaging clinical variables obtained based on data from the Breast Cancer Surveillance Consortium (BCSC). Methods: BCSC collected data on women undergoing mammography at several registries across the US. Our study cohort consisted of women who eventually got diagnosed with breast cancer. We considered various methods for model building, specifically, multinomial logistic regression, linear discriminant analysis, quadratic discriminant analysis, naïve Bayes, decision tree, bagging, boosting, random forest, and feed-forward neural network models. Multiple imputation via chained equations was used to address the missing data. Results: We analyzed data from 37,936 women aged 18–88 years at the time of mammogram recorded before their breast cancer diagnosis. The multinomial logistic regression model achieved the highest classification accuracy and a multi-class AUC of 0.704 Predictors in the final model are age at mammogram, race/ethnicity, age at first childbirth, menopausal status, hormone replacement therapy use, first-degree family history of breast cancer, BMI category, and history of breast biopsy. Conclusion: Although accuracy was moderate, the model reliably estimated BI-RADS breast density using routinely available predictors. By leveraging easily measurable factors, the model can be applied to estimate breast density for women without mammographic data, allowing the possibility of preliminary risk assessment.
Slides
Photos
Participants: 8