TruaceTracing the truth around AITuesday, September 15, 2026
The Index

What the evidence says.What the public feels.

Ranks distinct AI gain and problem claims from the published record. Scores reward impact, independent source strength, scale, confidence, and recency.

1,419 results
Show filters and sorting

AI gains · 788

73
GainHealth· Stable· Evidence: Moderate (1 source)

ChatGPT provided satisfactory answers to common patient hip arthroscopy questions, with half of responses graded A and another 30% graded B by fellowship-trained surgeons.

In a study published June 22, 2024, two hip preservation surgeons graded ChatGPT 3.5 answers to ten common hip arthroscopy questions drawn from patient education sites, using an A-to-D scale and readability scores FRES and FKGL.

Impact 30%69
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%89

Updated Jul 20, 2026 · TRV-2026-0412

73
GainHealth· Stable· Evidence: Moderate (1 source)

AI-enabled hierarchical medical system increased hypertension control target achievement from 68% to 82% and achieved health data accuracy exceeding 95% for chronic disease patients.

Researchers assessed an AI-driven hierarchical medical system for chronic disease management in China using 2024 National Health Commission monitoring data and 12,468 patient follow-up records from three provinces, structured around data collection, decision intervention, resource scheduling, and outcome feedback.

Impact 30%69
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%89

Updated Jul 20, 2026 · TRV-2026-0328

73
GainHealth· Stable· Evidence: Moderate (1 source)

In routine practice between May 2023 and May 2025, Brainomix e-CTA achieved 84% sensitivity and 95% specificity for LVO detection and reduced time to diagnostic conclusion for readers of different experience levels, with 94% sensitivity for ICA/proximal M1 occlusions.

Between May 2023 and May 2025, researchers retrospectively evaluated 531 multiphase CTA examinations from consecutive patients with suspected acute ischemic stroke at a single center, comparing Brainomix e-CTA automated LVO detection to expert neuroradiologist interpretation.

Impact 30%69
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%89

Updated Jul 20, 2026 · TRV-2026-0307

73
GainHealth· Stable· Evidence: Moderate (1 source)

Convolutional neural networks trained on four Scheimpflug-based corneal maps differentiated keratoconus from normal astigmatic eyes with up to 99.2% accuracy and AUC 1.00, with external validation retaining 97-98% accuracy.

By July 2026, a cross-sectional study at Al-Shifa Trust Eye Hospital in Pakistan developed four CNN models on 5602 Scheimpflug-derived corneal maps from 1411 eyes to distinguish keratoconus from normal eyes, reporting internal accuracies of 98.1% to 99.2% and AUCs up to 1.00, with external validation on 85 participants confirming 97.1% to 98.3% accuracy.

Impact 30%69
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%89

Updated Jul 18, 2026 · TRV-2026-0256

AI problems · 631

68
ProblemScience· Newly added· Evidence: Moderate (1 source)

AI models for craniofacial soft tissue prediction lack independent external validation, leaving generalizability beyond the study dataset unestablished.

By September 2026, researchers had built an AI-assisted framework using CT scans from 1972 individuals in northwestern India to jointly characterize craniofacial soft tissue thickness and subcutaneous fat thickness at 73 landmarks, analyzing sexual dimorphism, age variation, and bilateral symmetry.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%99

Updated Sep 10, 2026 · TRV-2026-1051

68
ProblemHealth· Newly added· Evidence: Moderate (1 source)

Half of included studies had identified risk of bias and clinical adoption remains limited, requiring multicenter prospective validation and interpretable frameworks before bedside use.

A PRISMA-compliant systematic review and meta-analysis of 8 studies with 5503 adult patients evaluated supervised machine learning models to predict all-cause mortality in infective endocarditis, a condition described as often fatal despite surgical and antibiotic advances. Five studies were pooled, showing strong discrimination for both in-hospital/early and 6-month mortality.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%99

Updated Sep 10, 2026 · TRV-2026-1049

68
ProblemHealth· Newly added· Evidence: Moderate (1 source)

Visual HRCT scoring remains reader-dependent and AI outputs lack prospective multicenter validation and protocol harmonization needed to serve as treatment-triggering biomarkers.

This peer-reviewed review summarizes recent AI and radiomics work in systemic sclerosis-associated ILD, the leading cause of disease-related mortality in systemic sclerosis, where visual HRCT scoring is reader-dependent. It reports that deep-learning UIP probability and whole-chest quantitative biomarkers stratify FVC decline and survival, and that AI-derived HRCT parameters correlate with DLCO/TLC and support pattern classification across CTD-ILD subtypes.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%99

Updated Sep 10, 2026 · TRV-2026-1048

68
ProblemHealth· Newly added· Evidence: Moderate (1 source)

Chatbots sometimes confused similar congenital anomalies and provided less detailed anatomical descriptions in complex tumor cases, requiring caution and verification by qualified professionals before clinical use.

On September 9, 2026, a peer-reviewed comparative study reported testing ChatGPT, Gemini, and Microsoft Copilot on 20 radiological cases split between congenital anomalies and tumors. Each system received the same questions and images and was scored for diagnostic accuracy and explanatory completeness, with ChatGPT scoring 19 correct, Gemini 17, and Copilot 16.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%99

Updated Sep 10, 2026 · TRV-2026-1047

Recomputed live from the record · Sep 15, 2026, 8:33 PM