Leveraging time-series electronic health records with large language models for chronic kidney disease diagnosis in primary care
Abstract: Chronic kidney disease (CKD) presents a growing public health challenge in China, exacerbated by low patient awareness and limited nephrology resources. This study evaluated the potential of large language models (LLMs) to support CKD diagnosis in primary care using time-series electronic health record (EHR) data. Longitudinal data were extracted from the EHR system of Weinan, China. Among 29 963 adults meeting inclusion criteria, 2300 participants were randomly sampled to generate clinical vignettes. CKD status…

"Bill on dialysis" by fireflythegreat is licensed under CC BY 2.0. To view a copy of this license, visit https://creativecommons.org/licenses/by/2.0/.
Chronic kidney disease (CKD) presents a growing public health challenge in China, exacerbated by low patient awareness and limited nephrology resources. This study evaluated the potential of large language models (LLMs) to support CKD diagnosis in primary care using time-series electronic health record (EHR) data.
Using the DeepSeek LLM with prompt engineering, we generated binary CKD diagnoses and probability scores based on EHR data. Diagnostic performance was evaluated using accuracy, sensitivity, specificity, F1 score, area under the curve (AUC) and the detection rate for early kidney injury and compared with traditional machine learning (ML) models.
- Chronic kidney disease (CKD) presents a growing public health challenge in China, exacerbated by low patient awareness and limited nephrology resources.
- This study evaluated the potential of large language models (LLMs) to support CKD diagnosis in primary care using time-series electronic health record (EHR) data.
- Longitudinal data were extracted from the EHR system of Weinan, China.
Using 1-month EHR data, the LLM achieved an accuracy of 87.0%, AUC of 0.921, F1 score of 0.579, sensitivity of 89.1%, specificity of 86.8%, and detection rate for early kidney injury of 42.2%.
The rundown
Longitudinal data were extracted from the EHR system of Weinan, China. Among 29 963 adults meeting inclusion criteria, 2300 participants were randomly sampled to generate clinical vignettes.
Sources
- Peer-reviewedBMJ Health & Care Informatics2026-09-18
How should this claim be treated?
ace
The debate