TruaceTracing the truth around AIWednesday, September 16, 2026
The Index

What the evidence says.What the public feels.

Ranks distinct AI gain and problem claims from the published record. Scores reward impact, independent source strength, scale, confidence, and recency.

1,439 results
Show filters and sorting

AI gains · 798

77
GainHealth· Stable· Evidence: Moderate (1 source)

A Random Forest model trained on linguistic, emotional, cognitive, behavioral and temporal features from Weibo posts predicted Self-Rating Anxiety Scale scores among consenting Chinese college students with R2 0.77 on the test set.

Researchers surveyed college students in China with the Self-Rating Anxiety Scale and, with informed consent, analyzed their public Weibo posts. Using multi-dimensional features, a Random Forest model predicted anxiety scores within the study sample, achieving the best test performance among four models tested.

Impact 30%49
Evidence 25%95
Scale 20%85
Confidence 15%87
Recency 10%87

Updated Jul 13, 2026 · TRV-2026-0121

75
GainHealth· Newly added· Evidence: Moderate (1 source)

A reliability-oriented hybrid framework combining a periocular ResNet18 and a facial ensemble of ResNet50, EfficientNet-B0 and DenseNet121 increased early ASD risk indication to 90% sensitivity in the periocular pathway, 87.1% sensitivity with 0.948 AUC in the facial pathway, and an analytically estimated 98.71% system

Researchers developed a hybrid deep learning system for early autism spectrum disorder risk indication that fuses a ResNet18 model trained on static periocular images with a multi-CNN facial ensemble of ResNet50, EfficientNet-B0 and DenseNet121, combining outputs through an OR-based reliability rule and using Grad-CAM to highlight decision regions.

Impact 30%69
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%100

Updated Sep 16, 2026 · TRV-2026-1111

75
GainHealth· Newly added· Evidence: Moderate (1 source)

Combining Aidoc AI with a human radiologist increased sensitivity for intracranial hemorrhage on non-contrast head CT to 96.0% while maintaining 99.4% specificity, detecting cases missed by radiologists alone.

In a retrospective study of 4027 consecutive non-contrast head CT examinations from an emergency hospital in southwest Sweden, researchers compared three commercial AI algorithms for intracranial hemorrhage detection against reports from two radiologists, using two-tier consensus adjudication as the reference standard for 385 positive or discrepant cases.

Impact 30%69
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%100

Updated Sep 16, 2026 · TRV-2026-1107

AI problems · 641

73
ProblemHealth· Stable· Evidence: Moderate (1 source)

Digital mental health tools are hampered by engagement challenges, industry setbacks, methodological critiques, and gaps in evidence and scaling that limit real-world applicability.

As of May 2025, this review in World Psychiatry examined how smartphone apps, virtual reality, and generative AI including large language models are being applied to mental health, evaluating evidence across well-being, depression, anxiety, schizophrenia, eating disorders and substance use, and outlining advances in digital phenotyping and generative outputs.

Impact 30%49
Evidence 25%95
Scale 20%60
Confidence 15%87
Recency 10%90

Updated Jul 24, 2026 · TRV-2026-0523

73
ProblemHealth· Stable· Evidence: Moderate (1 source)

Across 29 standardized clinical vignettes, all 21 tested LLMs failed differential diagnosis in over 80% of cases, indicating they have not achieved the reasoning needed for safe clinical deployment.

Researchers evaluated 21 off-the-shelf large language models, including GPT-5, Claude 4.5 Opus, Gemini 3.0 and Grok 4, on 29 standardized MSD Manual clinical vignettes representing 16,254 responses scored by medical students. Using the PrIME-LLM composite across differential diagnosis, diagnostic testing, final diagnosis, management, and miscellaneous reasoning, scores ranged from 0.64 to 0.78.

Impact 30%69
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%87

Updated Jul 13, 2026 · TRV-2026-0145

72
ProblemHealth· Stable· Evidence: Moderate (1 source)

Heterogeneous NDC, Multum, and RxCUI identifiers in real-world EHRs undermined semantic consistency, with over half of records needing string reconciliation and up to 57.4% requiring correction due to branded formulation omissions and indication- or route-based ATC ambiguities.

Researchers developed and tested an informatics framework to convert heterogeneous discharge medication identifiers from EHRs of adults 65 and older at Buffalo General Medical Center between 2020 and 2024 into standardized RxCUI ingredient and ATC class codes. Of 214,080 records, 53% were nonstandardized Multum IDs requiring string-based reconciliation, and the team measured mapping success and correction needs after deterministic crosswalks and expert validation.

Impact 30%63
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%97

Updated Sep 1, 2026 · TRV-2026-0957

Recomputed live from the record · Sep 16, 2026, 4:11 PM