TruaceTracing the truth around AIWednesday, September 16, 2026
The Index

What the evidence says.What the public feels.

Ranks distinct AI gain and problem claims from the published record. Scores reward impact, independent source strength, scale, confidence, and recency.

1,439 results
Show filters and sorting

AI gains · 798

68
GainScience· Stable· Evidence: Moderate (1 source)

Generative AI applied to supply chain and operations management can enhance decision-making and optimize processes across areas like demand forecasting and inventory management to improve efficiency, accuracy, and resilience.

Researchers developed a capability-based framework to analyze where artificial intelligence and generative AI fit into supply chain and operations management. Using capabilities like learning, perception, prediction, interaction, adaptation and reasoning, they mapped applications across 13 decision areas including demand forecasting, inventory management, supply chain design and risk management.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%93

Updated Aug 12, 2026 · TRV-2026-0739

68
GainHealth· Stable· Evidence: Moderate (1 source)

Guideline-specific and general-purpose LLMs produced responses broadly consistent with EAU erectile dysfunction recommendations, with Gemini 2.5 Pro and EAU Guidelines Bot achieving the highest composite scores.

On 2026-08-07, a peer-reviewed comparative study tested five AI systems including the EAU Guidelines Bot, ChatGPT-5, Gemini 2.5 Pro, Copilot - Smart GPT-5, and Perplexity Pro on 13 questions drawn from strongly recommended EAU erectile dysfunction statements. Three senior reviewers rated each answer for relevance, clarity, structure, clinical utility, and factual accuracy.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%93

Updated Aug 10, 2026 · TRV-2026-0727

68
GainCrime· Stable· Evidence: Moderate (1 source)

Frontier chatbots showed less concerning behavior in newer models and when early escalation interventions were applied during mental-health conversations.

On 2026-08-07, Nature Medicine published a clinically validated auditing framework called SIM-VAIL that simulates users with psychiatric vulnerabilities to test frontier chatbots including Claude, ChatGPT, Gemini, Grok and Llama. Across 810 multi-turn conversations with 30 simulated profiles and scoring on 13 risk dimensions, the study observed widespread concerning behavior that accumulated over turns.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%93

Updated Aug 10, 2026 · TRV-2026-0726

68
GainHealth· Stable· Evidence: Moderate (1 source)

AI augmentation of operational workflows under human oversight improves oncology trial feasibility and patient identification, with tools for enrollment screening and monitoring now implemented at select cancer centres.

On 2026-08-07, a Review in Nature Reviews Clinical Oncology described how AI enabled by electronic health record datasets and machine learning is being applied across pre-trial design, conduct, and post-trial inference in oncology. It reported that the most immediate evidence-supported uses are operational workflows under human oversight, including patient identification, eligibility assessment, data extraction, and trial monitoring, now implemented at select cancer centres.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%93

Updated Aug 10, 2026 · TRV-2026-0725

AI problems · 641

67
ProblemEducation· Stable· Evidence: Moderate (1 source)

Generative AI adoption in higher education created persistent risks to academic integrity, data privacy, equity, and responsible governance.

This March 2026 systematic and thematic review examined generative AI tools such as ChatGPT in higher education, analyzing 46 Web of Science documents and qualitatively synthesizing 27 peer-reviewed articles to map implementation trends.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%89

Updated Jul 20, 2026 · TRV-2026-0342

67
ProblemHealth· Stable· Evidence: Moderate (1 source)

In patients with LAD myocardial bridging, longer bridging length increased risk of abnormal FFRCT, with a larger effect in females, and females with isolated bridging had more pronounced distal hemodynamic compromise.

A retrospective study of 300 patients with left anterior descending artery myocardial bridging and 104 controls used an AI-based platform, Shukun-FFRCT, to obtain whole-vessel and segmental FFRCT values and relate them to bridging morphology and sex.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%89

Updated Jul 20, 2026 · TRV-2026-0331

67
ProblemHealth· Stable· Evidence: Moderate (1 source)

Use of the same ambient AI scribes was associated with increased length of notes and no change in physician productivity measured by billing metrics.

A rapid review published April 29 2025 synthesized 6 real-world studies of digital scribes using ambient listening and generative AI from 1450 screened records spanning academic health systems, community settings, and outpatient practices. Across observational, case report, cohort, and survey designs, authors reported decreased self-reported documentation times with associated increased length of notes.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%89

Updated Jul 20, 2026 · TRV-2026-0330

67
ProblemHealth· Stable· Evidence: Moderate (1 source)

Despite higher scores, the evaluated large language models have identified limitations that prohibit them from replacing human experts for triage in overcrowded emergency departments.

Researchers designed the Skyer benchmark to evaluate fifteen large language models on 55 realistic pediatric emergency department scenarios using a weighting system for over-triage and under-triage plus three repeat runs for consistency. By the publication date of July 11 2026, ChatGPT-4.5-preview and Gemini-2.5_05-06 had shown 77% and 74% accuracy with mean weights of 377.5 and 365 out of 550, compared to 64% and 253.5 for human experts.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%89

Updated Jul 20, 2026 · TRV-2026-0327

Recomputed live from the record · Sep 16, 2026, 1:54 PM