TruaceTracing the truth around AITuesday, September 15, 2026
The Index

What the evidence says.What the public feels.

Ranks distinct AI gain and problem claims from the published record. Scores reward impact, independent source strength, scale, confidence, and recency.

1,419 results
Show filters and sorting

AI gains · 788

68
GainHealth· Newly added· Evidence: Moderate (1 source)

XGBoost models trained on 5798 participants identified depression risk among older adults with chronic diseases across cognitive impairment levels with accuracy up to 0.767 and good calibration.

Researchers developed three XGBoost models to identify depression risk among older adults with chronic illnesses, stratified by cognitive impairment status, using 5798 participants from the Chinese Longitudinal Healthy Longevity Survey and SHAP for interpretability.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%99

Updated Sep 13, 2026 · TRV-2026-1070

68
GainHealth· Newly added· Evidence: Moderate (1 source)

Large language models including ChatGPT-o3, ChatGPT-5.2, Gemini 3 and DeepSeek provided generally acceptable clinical accuracy when answering 30 frequently asked patient questions about robotic-assisted total knee arthroplasty.

A September 2026 peer-reviewed study in The Knee compared four large language models on 30 common patient questions about robotic-assisted total knee arthroplasty, evaluating responses with DISCERN, QAMAI, a 5-point clinical accuracy scale, and PEMAT and Flesch-Kincaid readability measures.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%99

Updated Sep 13, 2026 · TRV-2026-1069

68
GainHealth· Newly added· Evidence: Moderate (1 source)

Across 17 studies totaling 243,324 trauma patients, the best-performing ML model showed a small pooled AUC advantage over logistic regression for mortality prediction.

A systematic review and meta-research appraisal examined 20 studies comparing machine learning and logistic regression for trauma mortality prediction, with 17 studies (243,324 patients) in primary synthesis. The pooled within-study AUC difference favoring the best ML model was 0.026 (95% CI 0.009-0.043), 0.017 in co-primary analysis of studies reporting CIs, with extreme heterogeneity and a prediction interval crossing zero.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%99

Updated Sep 13, 2026 · TRV-2026-1068

68
GainHealth· Newly added· Evidence: Moderate (1 source)

Deep learning models showed promising and often strong performance for predicting conversion from mild cognitive impairment to Alzheimer's disease, supporting early diagnosis and timely therapeutic intervention.

This PRISMA-guided systematic review examined 60 studies published between 2019 and February 2026 that used deep learning to classify Alzheimer's stages and predict conversion from mild cognitive impairment to Alzheimer's disease. It found cross-sectional designs predominant, CNNs dominant for neuroimaging, and growing use of RNNs and transformers for longitudinal data, with multimodal approaches in 24 studies.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%99

Updated Sep 13, 2026 · TRV-2026-1067

AI problems · 631

68
ProblemHealth· Stable· Evidence: Moderate (1 source)

AI facial image scoring, editing and curation systems converge on a narrow westernized phenotype and are linked to appearance dissatisfaction, perception drift, and Snapchat and Zoom dysmorphia presentations in plastic surgery patients.

This educational review from August 2026 examines how AI systems that score, edit, generate and curate facial images encode explicit quantitative definitions of attractiveness and deliver them through prediction algorithms, AR filters, generative imagery and surgical simulators. It finds independent models converge on a narrow, frequently westernized phenotype and summarizes evidence linking such imagery to appearance dissatisfaction and perception drift.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%94

Updated Aug 16, 2026 · TRV-2026-0780

68
ProblemHealth· Stable· Evidence: Moderate (1 source)

Because the dataset was split at the image level, correlated frames from the same examination may inflate performance, so results cannot be interpreted as patient-level generalization and require grouped reanalysis and external validation before clinical use.

A comparative study merged SEE-AI and Kvasir-Capsule into a 21-class capsule endoscopy image dataset and fine-tuned a Vision Transformer, DenseNet121, and ResNet50. On an independent test set of 8,696 frames, the transformer achieved 92.2% accuracy and 0.99 AUC, substantially higher than the two CNN baselines under the reported experimental conditions.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%94

Updated Aug 16, 2026 · TRV-2026-0779

68
ProblemHealth· Stable· Evidence: Moderate (1 source)

Deep learning OCT models trained on a single dataset often degrade across scanners and sites due to device-dependent speckle variability, limiting reliability in real-world screening.

Researchers developed NA-DyCNN, a lightweight noise-aware dynamic convolutional network for OCT-based retinal disease classification that explicitly models post-acquisition speckle variability during training to improve cross-scanner robustness.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%94

Updated Aug 15, 2026 · TRV-2026-0775

68
ProblemHealth· Stable· Evidence: Moderate (1 source)

Most AI-ECG studies in pediatric and congenital heart disease remain retrospective and single-center with limited external validation, facing challenges from small datasets and age-dependent ECG variation.

This review from August 2026 summarizes how artificial intelligence applied to standard electrocardiograms has been tested in pediatric and congenital heart disease. It reports that deep learning models have been shown to identify arrhythmias, ventricular dysfunction, and CHD, and are being extended to predict future risk and to analyze wearable and telemetry data.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%94

Updated Aug 15, 2026 · TRV-2026-0774

Recomputed live from the record · Sep 16, 2026, 3:11 AM