TRV-2026-0562Version 1 · Certified
Reason for this version
Certified into the record
Canonical text (the exact bytes fingerprinted)
TRUVACE RECORD VERSION record: TRV-2026-0562 version: 1 kind: certified reason: Certified into the record timestamp: 2026-07-25T06:09:21.427062Z status: published lens: g_space sector: health headline: Persona-Driven Data Augmentation for Disease Name Recognition Across Rare and General Disease Corpora: Comparative Evaluation Study dek: Medical information extraction requires automatically identifying disease names and related terms in text. This task, known as named entity recognition (NER), relies on expert-annotated data that are costly to produce and often available only in limited quantities. Data augmentation (DA) aims to expand available training data; however, standard techniques such as synonym replacement and back-translation may introduce inappropriate substitutions or fail to preserve entity-label alignment, which is critical for se… gain_title: Persona-driven document-level augmentation with multiple LLM personas increased BioBERT disease NER F1 over gold-standard-only training on both RareDis and NCBI disease datasets. problem_title: (none) trace_subject: (none) gain_reading: Persona-driven document-level augmentation with multiple LLM personas increased BioBERT disease NER F1 over gold-standard-only training on both RareDis and NCBI disease datasets. gain_evidence: Persona-driven DA improved NER performance over GS-only training in both datasets problem_reading: (none) problem_evidence: (none) quick_read: Researchers tested whether large language model rephrasing controlled by persona prompts and XML tags could expand scarce expert-annotated data for disease named entity recognition. They applied the method to RareDis, a low-resource rare disease corpus, and NCBI disease, a general disease benchmark, and compared BioBERT performance with and without augmented variants. The work matters because disease name recognition underpins medical information extraction, yet expert annotation is costly and limited. Results show controlled linguistic variation can lift F1 scores and reduce data requirements in some settings, but gains were not uniform across corpora and entity types, leaving open how well the approach transfers to other clinical NER tasks and real-world deployment. limitation: Benefit varied across datasets with more modest gains in the low-resource rare disease corpus, indicating dataset-dependent effectiveness. tag: Evidence-backed gain key_points: Framework used multiple personas varying in medical expertise, personality, tone, and narrative style with XML-tag constrained prompting to preserve annotated entity spans. | Evaluated on RareDis low-resource rare disease corpus and NCBI disease general benchmark using microaveraged entity-level precision, recall, and F1-score. | Best RareDis result came from low-fidelity persona subset at 73.35 F1 versus 71.22 baseline; best NCBI result from all-personas at 89.32 versus 87.82 baseline. | In low-resource experiments on NCBI disease, all-personas and high-fidelity settings exceeded 100% gold-standard performance using only 60% of training data. rundown: The study built a document-level augmentation pipeline where each persona rephrased training documents while aiming to preserve annotated entity spans, measured semantic fidelity with BERTScore and lexical diversity with BLEU-4, and grouped personas into high-, balanced-, and low-fidelity subsets. BioBERT models were fine-tuned under gold-standard only, synonym replacement, single-persona, curated subsets, and all-persona conditions; entity-level analysis showed improvements across RareDis categories particularly for symptom and reduced symptom-sign confusion. sources: - peer_reviewed | JMIR Medical Informatics | https://doi.org/10.2196/87831 | 2026-07-24 prev: 0000000000000000000000000000000000000000000000000000000000000000
- sha256
- 6d206d70429d7952200cec039e20bcde6f95db2c0f7f528d76e69a8276a3a292
- previous
- 0000000000000000000000000000000000000000000000000000000000000000
Verify this record
How to verify without trusting this page
Fetch the canonical text of any version from /api/record/TRV-2026-0562 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.
ace