Persona-Driven Data Augmentation for Disease Name Recognition Across Rare and General Disease Corpora: Comparative Evaluation Study
Medical information extraction requires automatically identifying disease names and related terms in text. This task, known as named entity recognition (NER), relies on expert-annotated data that are costly to produce and often available only in limited quantities. Data augmentation (DA) aims to expand available training data; however, standard techniques such as synonym replacement and back-translation may introduce inappropriate substitutions or fail to preserve entity-label alignment, which is critical for se…
Persona-driven document-level augmentation with multiple LLM personas increased BioBERT disease NER F1 over gold-standard-only training on both RareDis and NCBI disease datasets.
Benefit varied across datasets with more modest gains in the low-resource rare disease corpus, indicating dataset-dependent effectiveness.
Evidence
- Peer-reviewedJMIR Medical Informatics2026-07-24
How should this claim be treated?
Truvace Impact Record TRV-2026-0562, v1: “Persona-Driven Data Augmentation for Disease Name Recognition Across Rare and General Disease Corpora: Comparative Evaluation Study.” Truvace, 2026-07-25. /record/TRV-2026-0562 (accessed at citation time). sha256 6d206d70429d7952…
Calibration history
Every change to this record since certification, in the open. None yet — the reading has held since it entered the record.
Certified into the record
How to verify without trusting this page
Fetch the canonical text of any version from /api/record/TRV-2026-0562 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.
ace