The Health-Wealth Gradient in Labor Markets: Integrating Health, Insurance, and Social Metrics to Predict Employment Density
Labor market forecasting relies heavily on economic time-series data, often overlooking the “health–wealth” gradient that links population health to workforce participation. This study develops a machine learning framework integrating non-traditional health and social metrics to predict state-level employment density. Methods: We constructed a multi-source longitudinal dataset (2014–2024) by aggregating county-level Quarterly Census of Employment and Wages (QCEW) data with County Health Rankings to the state lev…

"Minimum Wage Event with US Secretary Perez at Boloco" by MDGovpics is licensed under CC BY 2.0. To view a copy of this license, visit https://creativecommons.org/licenses/by/2.0/.
Researchers developed a machine learning framework that combines economic employment data with non-traditional health and social metrics to forecast employment density at the state level. Using county-level QCEW data aggregated with County Health Rankings from 2014 to 2024 and a time-aware validation across the COVID-19 break, a tuned regularized XGBoost model reached Test R2 = 0.800, with a stacked Ridge ensemble at 0.827, and SHAP values were used for interpretability.
The result suggests population health indicators can add predictive signal to labor market forecasting beyond standard economic time series, which could inform workforce planning and policy analysis. What remains uncertain from this excerpt is whether the accuracy gain translates into better decisions, how the model performs for specific states or demographic groups, and what costs or risks arise from using health data for employment prediction.
- Constructed multi-source longitudinal dataset 2014-2024 aggregating county-level QCEW with County Health Rankings to state level
- Compared LASSO, Random Forest, and regularized XGBoost using time-aware split across COVID-19 structural break
- Used SHAP values for interpretability of tree model
Integrating County Health Rankings and social metrics with QCEW data in a regularized XGBoost framework improved out-of-sample prediction of state-level employment density.
The rundown
The authors built a 2014-2024 longitudinal dataset by aggregating county-level Quarterly Census of Employment and Wages with County Health Rankings to the state level and evaluated models with a time-aware split across the COVID-19 break. They compared LASSO, Random Forest, and regularized XGBoost, using SHAP values to interpret drivers, and reported a leakage-safe stacked Ridge ensemble as a robustness check. The work frames the approach as capturing the health-wealth gradient linking population health to workforce participation. The source text does not report deployment, adoption, or downstream labor impacts beyond predictive accuracy. No harms, costs, or failure modes are quantified in the supplied excerpt. The evidence is limited to out-of-sample R2 on historical state aggregates, not prospective use by employers or agencies. No limitation or uncertainty language is included in the excerpt to substantiate a boundary. The publication date is 2026-01-15, so results are historical by
that date. The excerpt does not provide state names, effect sizes for health variables, or implementation details beyond model family and validation strategy.
Sources
- Peer-reviewedComputation2026-01-15
How should this claim be treated?
ace
The debate