TruaceTracing the truth around AIFriday, September 11, 2026
Health·The Trace·Dual reading·Published 2026-09-07

Mistral-NeMo classification of CPT codes for femur- and knee-related orthopaedic operative notes

Source article: Artificial Intelligence Accurately Assists in Billing for Orthopaedic Lower Extremity Surgery: Performance of the Mistral-NeMo Language Model

Abstract: Purpose To elucidate Mistral-NeMo's proficiency as a novel artificial intelligence assistant to verify coding accuracy and improve efficiency in manual medical coding practices in orthopaedic surgery. Methods This study tested Mistral-NeMo on 1000 operative notes labeled with the Current Procedural Terminology (CPT) codes from 177 providers. In total, there were 46 unique CPT codes; the most common were 29881 (knee arthroscopy with meniscectomy for both medial and lateral menisci), 29880 (knee arthroscopy with m…

TRV-2026-1003Peer-reviewedPermanent record — cite & verify
Trace impact reading

Contested: both sides are scored from claims and sources, not community votes.

P 68The P score combines the specificity and measured human impact of the grounded problem claim with the strength of this Trace’s cited sources.G 70The G score combines the specificity and measured human impact of the grounded gain claim with the strength of this Trace’s cited sources.
Artificial Intelligence Accurately Assists in Billing for Orthopaedic Lower Extremity Surgery: Performance of the Mistral-NeMo Language Model

Physical therapy in a pool agitator (SC 495871), National Museum of Health and Medicine (3300120120) by National Museum of Health and Medicine. CC BY 2.0 · https://creativecommons.org/licenses/by/2.0

The quick read

A peer-reviewed study tested the Mistral-NeMo language model on 1000 operative notes from 177 providers to verify Current Procedural Terminology codes for lower-extremity orthopaedic surgery. When CPT billing descriptions were included in the prompt, the model correctly identified 90% of true codes and rejected 99.80% of incorrect codes.

Automated validation could reduce incidental coding errors and administrative load that diverts clinicians and staff from patient care, but the reported high accuracy was observed only with descriptions provided. The study's scope was limited to femur- and knee-related notes and 46 codes, leaving open how the model would perform on broader procedures or without contextual descriptions.

Main points
  • Study tested Mistral-NeMo on 1000 operative notes labeled with CPT codes from 177 providers covering 46 unique CPT codes.
  • Most common codes were 29881, 29880, and 29888 for knee arthroscopy meniscectomy and ACL repair.
  • Binary Yes/No trials and confidence score 0-100 trials used true codes and randomly selected incorrect codes as controls.
Gain

Mistral-NeMo verified orthopaedic lower-extremity billing by correctly identifying 90% of true CPT codes and rejecting 99.8% of incorrect codes when provided with billing descriptions.

Problem

Mistral-NeMo failed to classify CPT codes accurately when billing descriptions were not provided, showing dependence on contextual information.

The rundown

Researchers prompted Mistral-NeMo with 1000 operative notes from 177 providers, each paired with either the true CPT code or a randomly selected incorrect code as positive or negative controls.

The dataset included 46 unique CPT codes, with knee arthroscopy codes 29881, 29880, and 29888 most frequent, and prompts requested binary Yes/No or 0-100 confidence responses.

What this doesn’t fix

Performance depended on contextual information and was insignificant without billing descriptions, limiting use on operative notes alone.

Sources

Reader signal

How should this claim be treated?

The debate