Development and Validation of Human-AI Collaborative Workflow in SNOMED CT Mapping of Bilingual Clinical Text
Systematized Nomenclature of Medicine-Clinical Terminology (SNOMED CT) is the principal international standard for semantic interoperability of clinical information, but mapping free-text clinical narratives to SNOMED CT concepts remains labor-intensive. We developed a large language model agent system for mapping bilingual clinical text to SNOMED CT concepts and evaluated its effect on mapping accuracy and efficiency within a human-AI collaborative workflow. We designed a three-module agent system comprising tr…

In brief
At a tertiary academic hospital in South Korea, researchers developed an LLM agent system to map bilingual free-text clinical narratives to SNOMED CT and tested it in a human-AI collaborative workflow against human-only mapping by three health information managers on 2,261 segments.
The collaborative approach matters because SNOMED CT mapping is labor-intensive but critical for semantic interoperability; the study shows accuracy gains alongside a 53.9% time reduction, though it remains uncertain how well the results transfer beyond one hospital, one language pair, and the nine categories studied.
Main points
- Study used 2,261 de-identified clinical text segments across nine clinical categories collected at a tertiary academic hospital in South Korea.
- System comprised translation, abbreviation expansion, and vector-based retrieval components integrated with a pre-embedded SNOMED CT vector database.
- Three health information managers compared human-only, Agent-only, and Agent-assisted human mapping using hit rate, precision, recall, F1 at k=1 and 5, and R-precision.
- Modular architecture supports periodic vector database updates without retraining for bilingual terminology standardization.
The gain
A three-module LLM agent system for bilingual clinical text expanded valid SNOMED CT candidates for expert mappers, raising hit rate@1 and R-precision while cutting per-segment mapping time by about half.
The rundown
Researchers built a three-module agent with translation, abbreviation expansion, and vector retrieval linked to a pre-embedded SNOMED CT database, then tested it on 2,261 de-identified bilingual segments from nine clinical categories.
Three health information managers mapped the same segments under human-only, Agent-only, and Agent-assisted conditions, with performance measured by hit rate, precision, recall, F1 at k=1 and 5, and R-precision, plus time per segment.
Sources
- Peer-reviewedJournal of Medical Systems2026-10-02
ace
The debate