Interpretable Machine Learning for Population-Level Tooth Loss Prediction
Machine learning can support population-level severe tooth loss (STL; ≥6 missing teeth) risk stratification; however, a lack of calibration under domain shift, limited interpretability of conventional black-box models, and inadequate handling of complex survey designs constrain responsible public health interpretation and implementation. We implemented and evaluated an interpretable, survey-weighted Multiple Imputation by Chained Equations-Explainable Boosting Machine (MICE-EBM) framework for population-level ST…

"dental xray" by Inha Leex Hale is licensed under CC BY 2.0. To view a copy of this license, visit https://creativecommons.org/licenses/by/2.0/.
Researchers developed a TRIPOD+AI-compliant, survey-weighted MICE-EBM framework to predict severe tooth loss defined as six or more missing teeth using US representative data. The model was derived on BRFSS 2022 with 433,772 adults, temporally validated on BRFSS 2024 with 448,213 adults, and tested for cross-survey generalizability on NHANES 2015-2018 with 10,775 adults.
The framework offers intrinsic glass-box interpretability and calibrated same-survey performance that could inform dental public health planning, but its value is constrained by degraded discrimination and calibration when transferred across surveys with different outcome definitions. Uncertainty remains about international applicability and whether recalibration alone is sufficient for responsible implementation in new populations.
- Study used BRFSS 2022 N=433,772 for derivation, BRFSS 2024 N=448,213 for temporal validation, and NHANES 2015-2018 N=10,775 for cross-survey evaluation.
- Missing data handled with antileakage HistGradientBoosting-driven MICE pipeline to preserve multivariate epidemiological variance.
- 100-replicate locked-imputed-set bootstrap on BRFSS 2022 showed optimism-corrected AUC 0.860 and Brier 0.088, confirming robustness.
- Predefined isotonic recalibration improved NHANES holdout Brier from 0.192 to 0.136 after direct transfer showed poor raw calibration.
Survey-weighted Explainable Boosting Machine achieved strong temporal stability for severe tooth loss prediction on US BRFSS cohorts, supporting transparent population-level risk stratification.
The rundown
Researchers derived the model on BRFSS 2022 and validated temporally on BRFSS 2024, reporting AUC 0.865 and Brier 0.086 for derivation and AUC 0.863 and Brier 0.085 for temporal validation.
Cross-survey evaluation on NHANES 2015-2018 tested robustness under domain shift, where raw transfer produced AUC 0.754 and Brier 0.192, improving to Brier 0.136 after isotonic recalibration, versus 0.780 AUC for a black-box stacked meta-ensemble.
Sources
- Peer-reviewedJournal of Dental Research2026-07-30
How should this claim be treated?
ace
The debate