Development and validation of a generalizable M-protein screening model using routine laboratory indicators: a multicenter retrospective study
Abstract: Early detection of plasma cell disorders (PCDs) remains challenging due to limited accessibility of gold standard diagnostic methods. This study aimed to develop a simple M-protein screening model using routine laboratory indicators for clinical laboratories. A total of 5217 participants from three Chinese hospitals were enrolled. The derivation cohort (n = 3019) was randomly divided into training and internal validation cohorts. Two external validation cohorts (n = 1747 and n = 451) were included. M-protein pos…

"NABL Accredited Pathology Lab and Diagnostic Centre in Kharghar, Navi Mumbai - Full Body Checkup, CBC, Thyroid Blood Test and Home Blood Collection" by Goleisureintl is licensed under CC BY 4.0. To view a copy of this license, visit https://creativecommons.org/licenses/by/4.0/.
Researchers developed and validated an M-protein screening model using routine laboratory indicators from 5217 participants across three Chinese hospitals. They compared eight machine learning algorithms and selected a logistic regression model incorporating sex, age, total protein, albumin, albumin/globulin ratio, and hemoglobin, achieving an AUC of 0.843 in training and 0.843, 0.801, and 0.800 in internal and two external validations with a five-tier risk stratification.
The work matters because gold standard methods SPE combined with IFE have limited accessibility, and a routine-lab-based model could support early diagnosis of clinically relevant plasma cell disorders in resource-limited primary healthcare settings. What remains uncertain is how the model performs outside the three-hospital Chinese population and whether prospective use changes clinical workflows or patient outcomes.
- Study enrolled 5217 participants from three Chinese hospitals with derivation cohort of 3019 and two external validation cohorts of 1747 and 451.
- M-protein positivity was defined by SPE combined with IFE and models used routine indicators including sex, age, total protein, albumin, albumin/globulin ratio, and hemoglobin.
- Eight algorithms were tested and logistic regression was selected as optimal with internal and external AUCs of 0.843, 0.801, and 0.800 and a five-tier risk stratification based on predicted probabilities.
Logistic regression screening model using routine indicators like total protein and hemoglobin achieved AUC 0.843 and provides an accessible tool for early identification of individuals at high risk of M-protein in resource-limited primary care.
The rundown
The authors collected demographic data and routine laboratory blood parameters and built eight models using logistic regression, k-nearest neighbors, decision tree, random forest, AdaBoost, linear discriminant analysis, quadratic discriminant analysis, and multilayer perceptron. The final models incorporated sex, age, total protein, albumin, albumin/globulin ratio, and hemoglobin.
Performance was evaluated in a derivation cohort randomly split into training and internal validation, plus two external validation cohorts. The selected LR model used a five-tier risk stratification based on predicted probabilities: 15.0%, 15.0-40.0%, 40.0-70.0%, 70.0-90.0%, and 90.0%, with reported AUCs of 0.843 internally and 0.801 and 0.800 externally.
Sources
- Peer-reviewedClinica Chimica Acta2026-08-26
How should this claim be treated?
ace
The debate