The Index recomputed live from the record

What the evidence says.What the public feels.

The record holds 933 sourced gains and 770 sourced problems, averaging 68 and 66 on the index score. Readers have logged 22 public signals on the Pulse, which is kept apart and never counted as evidence.

933 gains
770 problems
Every sourced claim in the Index, one square each, shaded by the strength of its evidence. High Moderate Emerging

Ranks distinct AI gain and problem claims from the published record. Scores reward impact, independent source strength, scale, confidence, and recency.

1,703 results
Show filters and sorting

Download every matching row, not just this page:Export CSVExport JSON

AI gains · 933

165
HealthStableModerate evidence · 1 source

A logistic regression model combining grayscale ultrasound, contrast-enhanced ultrasound, and serum TPO-Ab improved differentiation of benign versus malignant thyroid nodules in Hashimoto's thyroiditis, reaching cross-validated AUC 0.849 with 77.3% sensitivity and 77.8% specificity.

Researchers retrospectively analyzed 600 patients with Hashimoto's thyroiditis and 650 pathology-confirmed thyroid nodules to test whether combining grayscale ultrasound, contrast-enhanced ultrasound, and serum anti-thyroid peroxidase antibody improves malignancy differentiation. They built a logistic regression model on patients with complete data and performed stratified 5-fold cross-validation, also comparing six machine learning classifiers.

Impact 30%
63
Evidence 25%
95
Scale 20%
35
Confidence 15%
87
Recency 10%
94

Updated Sep 9, 2026 · TRV-2026-1027

72Index score
166
HealthStableModerate evidence · 1 source

Our method outperforms current algorithms in ultrasound PCa data, achieving mean Dice similarity coefficient (DSC), Jaccard similarity coefficient (OMG), and accuracy (ACC) of 83.6 ± 3.1%, 71.8 ± 2.5%, and 83.5 ± 3.1%, respectively.

Accurate Ultrasound (US) prostate cancer (PCa) segmentation images hold significant value for organ interventional guidance and clinical disease diagnosis. However, this task still poses substantial challenges.

Impact 30%
63
Evidence 25%
95
Scale 20%
35
Confidence 15%
87
Recency 10%
93

Updated Sep 3, 2026 · TRV-2026-0969

72Index score
167
HealthStableModerate evidence · 1 source

A two-layered RxCUI ingredient and ATC framework harmonized 214,080 discharge medication records from older adults into standardized representations, achieving 100% initial mapping via deterministic crosswalks to support transportable managed care AI tools.

Researchers developed and tested an informatics framework to convert heterogeneous discharge medication identifiers from EHRs of adults 65 and older at Buffalo General Medical Center between 2020 and 2024 into standardized RxCUI ingredient and ATC class codes. Of 214,080 records, 53% were nonstandardized Multum IDs requiring string-based reconciliation, and the team measured mapping success and correction needs after deterministic crosswalks and expert validation.

Impact 30%
63
Evidence 25%
95
Scale 20%
35
Confidence 15%
87
Recency 10%
92

Updated Sep 1, 2026 · TRV-2026-0957

72Index score
168
HealthStableModerate evidence · 1 source

A Random Forest model using eight routinely available variables predicted subsequent vasopressor need after two fluid boluses in pediatric suspected sepsis with AUROC 0.827 and stratified patients into four tiers with a 6.6% to 63.6% gradient.

Researchers developed a machine-learning risk stratification tool using routine EHR data from five pediatric emergency departments to predict need for vasoactive medication after two-bolus fluid resuscitation in suspected sepsis. Among 341 children meeting analytic criteria, 25.8% received vasopressors, and a Random Forest model achieved AUROC 0.827 and AUPRC 0.661 with four risk tiers.

Impact 30%
63
Evidence 25%
95
Scale 20%
35
Confidence 15%
87
Recency 10%
92

Updated Aug 31, 2026 · TRV-2026-0934

72Index score

AI problems · 770

165
ScienceNewly addedModerate evidence · 1 source

Observed high agreement may reflect a restricted score range, and the study did not validate LLM scores against human peer review.

In April 2026, researchers tested ChatGPT and Claude as peer reviewers by having each model score 50 general internal medicine abstracts twice under four fictional author identities, producing 800 evaluations of quality, novelty, and acceptance on a 0-10 scale.

Impact 30%
49
Evidence 25%
95
Scale 20%
35
Confidence 15%
87
Recency 10%
99

Updated Oct 4, 2026 · TRV-2026-1274

68Index score
166
PolicyNewly addedModerate evidence · 1 source

In pediatric care, LLMs often underperform relative to adult specialties and risk exacerbating health disparities due to biased or nonrepresentative training data.

On October 3 2026, Pediatrics published a policy statement on generative AI tools including large language models in pediatric health care. It describes potential uses in decision support, documentation, and education, while noting limited real-world validation and persistent concerns about accuracy, bias, and reliability.

Impact 30%
49
Evidence 25%
95
Scale 20%
35
Confidence 15%
87
Recency 10%
99

Updated Oct 4, 2026 · TRV-2026-1273

68Index score
167
HealthNewly addedModerate evidence · 1 source

Despite preserved discrimination, the algorithm showed calibration failure with compressed scores, low positive predictive value and low agreement, indicating distributional shift in this low- and middle-income country setting.

From 5 to 20 June 2026, researchers prospectively tested a publicly available DenseNet-121 model from TorchXRayVision in shadow mode on 826 consecutive chest radiographs from health assessment applicants at Patan Hospital in Lalitpur, Nepal, comparing AI scores to blinded single-reader radiologist classifications.

Impact 30%
49
Evidence 25%
95
Scale 20%
35
Confidence 15%
87
Recency 10%
98

Updated Oct 3, 2026 · TRV-2026-1265

68Index score
168
HealthNewly addedModerate evidence · 1 source

AI-integrated intraoral mobile photographic models showed high variability in specificity and have not yet been validated in real-world scenarios, requiring refinement of algorithms and standardization of imaging.

By October 2026, an umbrella review in International Journal of Dentistry synthesized prior systematic reviews on AI-integrated intraoral mobile photographs for screening oral potentially malignant disorders. It reported pooled sensitivity of 90% and specificity of 89% across included reviews, concluding the approach was effective as a noninvasive, cost-effective alternative to conventional invasive diagnostics.

Impact 30%
49
Evidence 25%
95
Scale 20%
35
Confidence 15%
87
Recency 10%
98

Updated Oct 3, 2026 · TRV-2026-1261

68Index score

Recomputed live from the record · Oct 11, 2026, 8:30 PM