The Index recomputed live from the record

What the evidence says.What the public feels.

The record holds 932 sourced gains and 770 sourced problems, averaging 68 and 66 on the index score. Readers have logged 22 public signals on the Pulse, which is kept apart and never counted as evidence.

932 gains
770 problems
Every sourced claim in the Index, one square each, shaded by the strength of its evidence. High Moderate Emerging

Ranks distinct AI gain and problem claims from the published record. Scores reward impact, independent source strength, scale, confidence, and recency.

1,702 results
Show filters and sorting

Download every matching row, not just this page:Export CSVExport JSON

AI gains · 932

85
HealthNewly addedModerate evidence · 1 source

ChatGPT-4o answered expert-validated true/false questions on vertical root fractures and tooth cracks with 86.1% accuracy, outperforming ChatGPT-3.5 and Gemini in a 5,400-response repeated test.

A peer-reviewed study in Odontology compared ChatGPT-3.5, ChatGPT-4o, and Google Gemini on 60 expert-validated true/false questions about longitudinal tooth fractures, querying each model three times daily for 10 days for 5,400 total responses against expert reference answers.

Impact 30%
69
Evidence 25%
95
Scale 20%
35
Confidence 15%
87
Recency 10%
99

Updated Oct 7, 2026 · TRV-2026-1305

74Index score
86
HealthNewly addedModerate evidence · 1 source

Automated extraction of psychological and subjective variables from unstructured forensic certificates was feasible with high specificity, with 24 of 35 variables meeting an 85% reliability threshold under at least one pipeline.

Researchers tested two automated pipelines on 110 randomly selected 2021 forensic medical certificates from 12 units of the French ORFeAD network, comparing a rule-based system using segmentation, lexicons and negation detection to a locally served Llama 3 8B model via Ollama, against physician coding of 35 binary variables.

Impact 30%
69
Evidence 25%
95
Scale 20%
35
Confidence 15%
87
Recency 10%
99

Updated Oct 5, 2026 · TRV-2026-1282

74Index score
87
HealthNewly addedModerate evidence · 1 source

AI-integrated intraoral mobile photographs achieved 90% pooled sensitivity and 89% specificity for noninvasive early detection of oral potentially malignant disorders, offering an accessible screening alternative in low-resource settings.

By October 2026, an umbrella review in International Journal of Dentistry synthesized prior systematic reviews on AI-integrated intraoral mobile photographs for screening oral potentially malignant disorders. It reported pooled sensitivity of 90% and specificity of 89% across included reviews, concluding the approach was effective as a noninvasive, cost-effective alternative to conventional invasive diagnostics.

Impact 30%
69
Evidence 25%
95
Scale 20%
35
Confidence 15%
87
Recency 10%
98

Updated Oct 3, 2026 · TRV-2026-1261

74Index score
88
HealthNewly addedModerate evidence · 1 source

A machine learning pipeline applied to urine SERS spectra improved multiclass diagnosis of IC and OAB versus healthy controls, achieving up to 92% accuracy with interpretable pathologic Raman features.

On October 2, 2026, a peer-reviewed study in ACS Sensors reported a urine-based diagnostic pipeline for interstitial cystitis and overactive bladder. Samples from 117 healthy controls, 19 IC patients, and 45 OAB patients were analyzed with Au-ZnO nanorod SERS chips, and spectra were classified using PCA-PLS-DA, PCA-LDA, XGBoost, and LightGBM, reaching up to 92% accuracy for three-way discrimination.

Impact 30%
69
Evidence 25%
95
Scale 20%
35
Confidence 15%
87
Recency 10%
98

Updated Oct 3, 2026 · TRV-2026-1256

74Index score

AI problems · 770

85
HealthStableModerate evidence · 1 source

Item Response Theory produced an unacceptably low 68% sensitivity for the total score at threshold 24, and logistic regression yielded only 16%-60% sensitivity for ADHD status in the same driver sample.

A 2026 peer-reviewed study validated the Persian Conners' Adult ADHD Rating Scale Short Version in 298 male taxi drivers in Iran, mean age 36.8, to establish occupational screening thresholds. Using a 198/100 train-test split, the authors compared ROC, item response theory, logistic regression and Random Forest approaches for cutoff selection.

Impact 30%
69
Evidence 25%
95
Scale 20%
35
Confidence 15%
87
Recency 10%
87

Updated Aug 3, 2026 · TRV-2026-0633

73Index score
87
HealthStableModerate evidence · 1 source

Across 29 standardized clinical vignettes, all 21 tested LLMs failed differential diagnosis in over 80% of cases, indicating they have not achieved the reasoning needed for safe clinical deployment.

Researchers evaluated 21 off-the-shelf large language models, including GPT-5, Claude 4.5 Opus, Gemini 3.0 and Grok 4, on 29 standardized MSD Manual clinical vignettes representing 16,254 responses scored by medical students. Using the PrIME-LLM composite across differential diagnosis, diagnostic testing, final diagnosis, management, and miscellaneous reasoning, scores ranged from 0.64 to 0.78.

Impact 30%
69
Evidence 25%
95
Scale 20%
35
Confidence 15%
87
Recency 10%
83

Updated Jul 13, 2026 · TRV-2026-0145

73Index score

Recomputed live from the record · Oct 11, 2026, 9:54 AM