TRV-2026-1282Version 1 · Certified

Written 2026-10-05 06:55:04 UTC · current record

Reason for this version

Certified into the record

Canonical text (the exact bytes fingerprinted)

TRUVACE RECORD VERSION
record: TRV-2026-1282
version: 1
kind: certified
reason: Certified into the record
timestamp: 2026-10-05T06:55:04.206099Z
status: published
lens: trace
sector: health
headline: Rule-Based Versus Generative Extraction of Psychological Symptoms From Forensic Medical Certificates: A Feasibility Study in the French ORFeAD Network
dek: Forensic medical certificates for victims of interpersonal violence describe psychological symptoms in unstructured prose, rarely explicit and requiring inference. To assess the feasibility of two extraction pipelines-rule-based and a locally served generative large language model-in a French multicentric forensic network, and to identify variables usable without human verification. 110 randomly selected 2021 certificates from 12 ORFeAD units were processed by both pipelines: rule-based (segmentation, lexicons,…
gain_title: Automated extraction of psychological and subjective variables from unstructured forensic certificates was feasible with high specificity, with 24 of 35 variables meeting an 85% reliability threshold under at least one pipeline.
problem_title: Sensitivity was heterogeneous and low for low-prevalence symptoms, 11 of 35 variables failed the reliability threshold under either pipeline, and neither pipeline supports individual-level decisions, with the generative model incurring far higher compute cost.
trace_subject: automated extraction of psychological symptoms from French ORFeAD forensic medical certificates for interpersonal violence victims
gain_reading: Automated extraction of psychological and subjective variables from unstructured forensic certificates was feasible with high specificity, with 24 of 35 variables meeting an 85% reliability threshold under at least one pipeline.
gain_evidence: Specificity was high for most variables | 24 of 35 variables met the threshold under at least one pipeline
problem_reading: Sensitivity was heterogeneous and low for low-prevalence symptoms, 11 of 35 variables failed the reliability threshold under either pipeline, and neither pipeline supports individual-level decisions, with the generative model incurring far higher compute cost.
problem_evidence: Sensitivity was heterogeneous and low for low-prevalence symptoms | neither pipeline supports individual-level decisions
quick_read: Researchers tested two automated pipelines on 110 randomly selected 2021 forensic medical certificates from 12 units of the French ORFeAD network, comparing a rule-based system using segmentation, lexicons and negation detection to a locally served Llama 3 8B model via Ollama, against physician coding of 35 binary variables.

The comparison matters because forensic certificates contain psychological symptoms in unstructured prose that requires inference, and reliable extraction could enable multicentric research on victim outcomes; however the feasibility results show uneven sensitivity, higher energy and GPU requirements for the generative approach, and explicit limits against using either pipeline for individual-level decisions without per-variable qualification.
limitation: Feasibility design with 110 certificates from 2021, no inferential testing, and out-of-the-box LLM settings limits generalizability and requires per-variable qualification before research use.
tag: Dual reading
key_points: 110 randomly selected 2021 certificates from 12 ORFeAD units were processed by both pipelines and coded by two forensic physicians with third-party arbitration. | Rule-based pipeline used segmentation, lexicons, regular expressions, negation detection; generative pipeline used Llama 3 8B, 4-bit, Ollama, near-default settings. | 35 binary variables, 23 psychological or subjective, were evaluated for sensitivity, specificity and accuracy with 85% reliability defined as usable.
rundown: The study compared a lexicon and regex rule-based system requiring 0.1-0.15 CPU-seconds per document against a locally served Llama 3 8B 4-bit model requiring 20-60 GPU-seconds, finding the local model did not clearly outperform at three orders of magnitude greater energy cost.

Authors coded 35 binary variables including 23 psychological or subjective ones described in unstructured prose that is rarely explicit and requiring inference, and concluded per-variable qualification is needed before any research use.
sources:
- peer_reviewed | Behavioral Sciences & the Law | https://doi.org/10.1002/bsl.70097 | 2026-10-03
prev: 0000000000000000000000000000000000000000000000000000000000000000
sha256
9b3945a0efaa9fce443df3034532bd9b0c4e3837c5efe24dc90272f56caf5cb4
previous
0000000000000000000000000000000000000000000000000000000000000000
Verify this record
How to verify without trusting this page

Fetch the canonical text of any version from /api/record/TRV-2026-1282 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.