TruaceTracing the truth around AITuesday, August 25, 2026
TRV-2026-0758Version 1 · Certified

Written 2026-08-14 06:23:02 UTC · current record

Reason for this version

Certified into the record

Canonical text (the exact bytes fingerprinted)

TRUVACE RECORD VERSION
record: TRV-2026-0758
version: 1
kind: certified
reason: Certified into the record
timestamp: 2026-08-14T06:23:02.258301Z
status: published
lens: p_space
sector: health
headline: Identifying the presence of disc herniations in lumbar spine MRI using Gemini 3.1 Pro
dek: Purpose Lumbar disc herniation is associated with substantial morbidity, including low back pain, radicular leg pain (sciatica), sensory disturbance, and motor deficit. Magnetic resonance imaging (MRI) is central to confirming the diagnosis in symptomatic patients and to planning surgical or interventional management. Recent advances in artificial intelligence (AI) raise the possibility of automating aspects of image interpretation to improve consistency and reduce radiologist workload. This study evaluates a ge…
gain_title: (none)
problem_title: When used zero-shot to identify lumbar disc herniations on sagittal MRI, Gemini 3.1 Pro produced low specificity and a substantial false-positive burden, with T1+T2 input performing worse than T1-only.
trace_subject: (none)
gain_reading: (none)
gain_evidence: (none)
problem_reading: When used zero-shot to identify lumbar disc herniations on sagittal MRI, Gemini 3.1 Pro produced low specificity and a substantial false-positive burden, with T1+T2 input performing worse than T1-only.
problem_evidence: false-positive burden was substantial, equivalent to 48.9 unnecessary reviews per 100 negative scans with T1 + T2
quick_read: Researchers tested Gemini 3.1 Pro in a zero-shot setting to detect lumbar disc herniations on sagittal MRI from 119 SPIDER cases (26% prevalence). Using only the mid-sagittal slice and a forced binary prompt, T1-only achieved 70% accuracy with 58% sensitivity and 74% specificity, while paired T1+T2 achieved 58% accuracy with 77% sensitivity and 51% specificity.

The findings matter because low specificity translates directly into unnecessary reviews and potential over-diagnosis in back-pain pathways, and the model fell below the ~80% specificity typically required for triage tools. Uncertainty remains about whether volumetric input, task-specific fine-tuning, or calibration could improve performance, as the current work was explicitly exploratory.
limitation: Evaluation was limited to zero-shot use of a generalist model on single mid-sagittal slices from a single paired cohort of 119 cases, presented as exploratory proof-of-concept rather than a validated clinical tool.
tag: Evidence-backed problem
key_points: Evaluated Gemini 3.1 Pro zero-shot on 119 cases from SPIDER public multi-center dataset with 31 herniation-positive and 88 negative. | T1-only input yielded sensitivity 0.58, specificity 0.74, accuracy 0.70; T1+T2 yielded sensitivity 0.77, specificity 0.51, accuracy 0.58. | Adding T2 reduced accuracy significantly by exact McNemar p = 0.044, driven by near-doubling of false positives from 23 to 43. | Authors concluded model is not currently suitable as triage or screening aid due to specificity below clinical threshold.
rundown: The study used SPIDER, a public multi-center dataset of sagittal T1- and T2-weighted lumbar MRI from patients with low back pain with expert level-by-level labels, testing the same 119 cases under T1-only and paired T1+T2 conditions using a standardized binary prompt.

Performance was reported with 95% confidence intervals and paired accuracy compared via exact McNemar test, with a supplementary evaluation using 3-slice image stacks to test 3-D processing ability.
sources:
- peer_reviewed | European Spine Journal | https://doi.org/10.1007/s00586-026-10285-9 | 2026-08-12
prev: 0000000000000000000000000000000000000000000000000000000000000000
sha256
6eb6a2cb0787f7b2f59934a65fe98fcf39f086b806ea57ffa95bd867e1e21ded
previous
0000000000000000000000000000000000000000000000000000000000000000
Verify this record
How to verify without trusting this page

Fetch the canonical text of any version from /api/record/TRV-2026-0758 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.