AI system for self-harm detection in emergency department triage notes
Source article: Temporal and cross-site validation of an AI system for self-harm detection
Abstract: Adequate self-harm surveillance is a key part of suicide prevention. Our previous research demonstrated that an artificial intelligence (AI)-based system could effectively detect self-harm in emergency department triage notes. However, the system was developed using data from a single hospital, raising concerns about its generalisability. Here, we aim to validate the system prospectively and externally to better understand its portability across hospitals. We leveraged emergency department data from two Australi…
Contested: both sides are scored from claims and sources, not community votes.
Hospital Universitari Doctor Peset, València 04 by 19Tarrestnom65. CC BY-SA 4.0 · https://creativecommons.org/licenses/by-sa/4.0
Researchers validated a previously developed AI system that detects self-harm in emergency department triage notes using extensive text normalisation and 1931 features. They tested it prospectively on 329,655 notes from the original major metropolitan hospital in Melbourne and externally on 316,877 notes from a regional hospital 150 km away covering 2012-2021.
The system maintained stable performance at the development site with PR AUC 0.84, supporting its epidemiological utility for self-harm surveillance, but declined to 0.78 at the regional site. The drop highlights portability challenges due to linguistic domain shift and differing presentation patterns, such as more medication ingestion cases regionally, leaving uncertainty about adaptation needs for broader deployment.
- System developed on 2012-2017 data from a major metropolitan hospital in Melbourne using 1931 selected features and extensive text normalisation.
- Prospective validation used 329,655 triage notes from same hospital over following four years; external validation used 316,877 notes from 2012-2021 from regional hospital 150 km outside Melbourne.
- Test set PR AUC was 0.84, 95% CI [0.82, 0.86], remaining stable prospectively at 0.84, 95% CI [0.83, 0.85].
- Regional performance declined to PR AUC 0.78, 95% CI [0.77, 0.79], attributed to linguistic domain shift and higher likelihood of medication ingestion presentations.
An AI system combining text normalisation with 1931 features maintained stable self-harm detection in prospective validation at its development metropolitan hospital, achieving PR AUC 0.84 over 329,655 triage notes in the following four years.
When applied to a regional hospital 150 km outside Melbourne, the same AI system's ability to distinguish self-harm cases declined to PR AUC 0.78, with instability linked to linguistic domain shift and different self-harm presentations.
The rundown
The study used manually annotated free-text triage notes from two Australian hospitals to test portability. Development data spanned 2012-2017 at the metropolitan site; prospective validation covered the next four years with 329,655 notes, while external validation used 316,877 notes from 2012-2021 at the regional site.
Text normalisation was equally effective across datasets, but the classification model faced domain shift. Authors noted regional self-harm presentations are more likely to involve medication ingestion, contributing to the observed PR AUC drop from 0.84 to 0.78.
System was developed from a single metropolitan hospital, limiting generalisability, and regional differences in presentation and language reduced stability.
Sources
- Peer-reviewedPLOS Digital Health2026-09-11
How should this claim be treated?
ace
The debate