TruaceTracing the truth around AIMonday, September 14, 2026
TRV-2026-1072Version 1 · Certified

Written 2026-09-13 06:57:05 UTC · current record

Reason for this version

Certified into the record

Canonical text (the exact bytes fingerprinted)

TRUVACE RECORD VERSION
record: TRV-2026-1072
version: 1
kind: certified
reason: Certified into the record
timestamp: 2026-09-13T06:57:05.865530Z
status: published
lens: trace
sector: health
headline: Temporal and cross-site validation of an AI system for self-harm detection
dek: Adequate self-harm surveillance is a key part of suicide prevention. Our previous research demonstrated that an artificial intelligence (AI)-based system could effectively detect self-harm in emergency department triage notes. However, the system was developed using data from a single hospital, raising concerns about its generalisability. Here, we aim to validate the system prospectively and externally to better understand its portability across hospitals. We leveraged emergency department data from two Australi…
gain_title: An AI system combining text normalisation with 1931 features maintained stable self-harm detection in prospective validation at its development metropolitan hospital, achieving PR AUC 0.84 over 329,655 triage notes in the following four years.
problem_title: When applied to a regional hospital 150 km outside Melbourne, the same AI system's ability to distinguish self-harm cases declined to PR AUC 0.78, with instability linked to linguistic domain shift and different self-harm presentations.
trace_subject: AI system for self-harm detection in emergency department triage notes
gain_reading: An AI system combining text normalisation with 1931 features maintained stable self-harm detection in prospective validation at its development metropolitan hospital, achieving PR AUC 0.84 over 329,655 triage notes in the following four years.
gain_evidence: This performance remained stable at the development site with a PR AUC of 0.84, 95% CI [0.83, 0.85]. | At the metropolitan hospital, the AI system for self-harm detection maintained its epidemiological utility.
problem_reading: When applied to a regional hospital 150 km outside Melbourne, the same AI system's ability to distinguish self-harm cases declined to PR AUC 0.78, with instability linked to linguistic domain shift and different self-harm presentations.
problem_evidence: When applied in the regional context, the model's ability to distinguish self-harm cases declined, resulting in an overall PR AUC of 0.78, 95% CI [0.77, 0.79]. | performance was unstable primarily due to linguistic domain shift and differences in self-harm presentations.
quick_read: Researchers validated a previously developed AI system that detects self-harm in emergency department triage notes using extensive text normalisation and 1931 features. They tested it prospectively on 329,655 notes from the original major metropolitan hospital in Melbourne and externally on 316,877 notes from a regional hospital 150 km away covering 2012-2021.

The system maintained stable performance at the development site with PR AUC 0.84, supporting its epidemiological utility for self-harm surveillance, but declined to 0.78 at the regional site. The drop highlights portability challenges due to linguistic domain shift and differing presentation patterns, such as more medication ingestion cases regionally, leaving uncertainty about adaptation needs for broader deployment.
limitation: System was developed from a single metropolitan hospital, limiting generalisability, and regional differences in presentation and language reduced stability.
tag: Dual reading
key_points: System developed on 2012-2017 data from a major metropolitan hospital in Melbourne using 1931 selected features and extensive text normalisation. | Prospective validation used 329,655 triage notes from same hospital over following four years; external validation used 316,877 notes from 2012-2021 from regional hospital 150 km outside Melbourne. | Test set PR AUC was 0.84, 95% CI [0.82, 0.86], remaining stable prospectively at 0.84, 95% CI [0.83, 0.85]. | Regional performance declined to PR AUC 0.78, 95% CI [0.77, 0.79], attributed to linguistic domain shift and higher likelihood of medication ingestion presentations.
rundown: The study used manually annotated free-text triage notes from two Australian hospitals to test portability. Development data spanned 2012-2017 at the metropolitan site; prospective validation covered the next four years with 329,655 notes, while external validation used 316,877 notes from 2012-2021 at the regional site.

Text normalisation was equally effective across datasets, but the classification model faced domain shift. Authors noted regional self-harm presentations are more likely to involve medication ingestion, contributing to the observed PR AUC drop from 0.84 to 0.78.
sources:
- peer_reviewed | PLOS Digital Health | https://doi.org/10.1371/journal.pdig.0001667 | 2026-09-11
prev: 0000000000000000000000000000000000000000000000000000000000000000
sha256
0f67e1fd3ba1457176b839e8e0aa0f23e129c15e8f4b9eab309d6265cd82325d
previous
0000000000000000000000000000000000000000000000000000000000000000
Verify this record
How to verify without trusting this page

Fetch the canonical text of any version from /api/record/TRV-2026-1072 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.