Improving 10-year cardiovascular disease risk prediction using automated machine learning
Abstract: Aims To develop a cardiovascular disease (CVD) risk prediction model with improved accuracy and interpretability by integrating diverse risk factors and applying Automated Machine Learning (AutoML), thereby enhancing clinical utility over conventional models. Methods This is a prospective cohort study. Data were obtained from the Multi-Ethnic Study of Atherosclerosis (MESA), including baseline and fifth follow-up visits, comprising 4713 participants. Exercise and dietary data were harmonized via Metabolic Equiva…
The Framingham study : an epidemiological investigation of cardiovascular disease sec.28 1973 by Kannel, William B., 1923-2011 Gordon, Tavia National Heart Institute (U.S.). Public domain
In a prospective cohort analysis of 4713 participants from the Multi-Ethnic Study of Atherosclerosis, researchers developed a 10-year CVD risk model integrating clinical, lifestyle, and cognitive measures. After selecting 21 predictors, they trained logistic regression, traditional machine learning models, and H2O AutoML, with AutoML achieving the highest performance at AUC 0.882 and accuracy 0.864 by the September 2026 publication date.
The result matters because improved discrimination and interpretability via SHAP could help clinicians stratify risk and target early prevention, potentially reducing population CVD burden. What remains uncertain is whether the observed performance generalizes beyond MESA and whether the highlighted Digit Symbol Score association represents a causal risk marker or a correlated indicator requiring further validation.
- Prospective cohort study used data from the Multi-Ethnic Study of Atherosclerosis including baseline and fifth follow-up visits with 4713 participants.
- Exercise and dietary data were harmonized via Metabolic Equivalent of Task (MET) and Healthy Eating Index-2015 (HEI-2015), and predictor selection used Boruta algorithm alongside Random Forest error rate cross-validation.
- SHAP analysis revealed relative importance of predictors, with age, TC and Digit Symbol Score (DSS) ranking highest, with DSS noted as a candidate risk marker.
H2O AutoML trained on 4713 MESA participants using 21 selected predictors achieved higher discrimination than logistic regression and traditional ML models, reaching AUC 0.882 and accuracy 0.864, intended as a practical tool for clinician risk stratification.
The rundown
Researchers used MESA data from 4713 participants across baseline and fifth follow-up visits, harmonizing lifestyle variables with MET and HEI-2015, selecting 21 predictors via Boruta and Random Forest cross-validation, and comparing logistic regression, four traditional machine learning algorithms, and H2O AutoML.
The best model was interpreted with SHapley Additive exPlanations, which ranked age, Total Cholesterol, and Digit Symbol Score as most important, supporting the authors' conclusion of superior discrimination and calibration for personalized prevention.
Sources
- Peer-reviewedInternational Journal of Cardiology2026-09-05
How should this claim be treated?
ace
The debate