TruaceTracing the truth around AITuesday, September 15, 2026
The Index

What the evidence says.What the public feels.

Ranks distinct AI gain and problem claims from the published record. Scores reward impact, independent source strength, scale, confidence, and recency.

1,418 results
Show filters and sorting

AI gains · 787

68
GainHealth· Stable· Evidence: Moderate (1 source)

The CODEX Action Incubator at UCSF convened 30 stakeholders and reached consensus on two priority metrics of AI scribe usage by primary care physicians to assess diagnostic excellence, including timely follow-up of abnormal breast and colorectal cancer screening results.

On September 2025, the CODEX Action Incubator at UCSF brought together 30 stakeholders from health systems, patient advocacy, industry and policy to address how to measure AI's effect on diagnostic excellence. Participants focused on AI scribes and, via a modified Delphi process, narrowed 17 candidate measures to two priority metrics tied to primary care physician usage rates: timely follow-up of abnormal breast and colorectal cancer screening results and patient-reported diagnostic experience.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%94

Updated Aug 14, 2026 · TRV-2026-0751

68
GainScience· Stable· Evidence: Moderate (1 source)

Generative AI applied to supply chain and operations management can enhance decision-making and optimize processes across areas like demand forecasting and inventory management to improve efficiency, accuracy, and resilience.

Researchers developed a capability-based framework to analyze where artificial intelligence and generative AI fit into supply chain and operations management. Using capabilities like learning, perception, prediction, interaction, adaptation and reasoning, they mapped applications across 13 decision areas including demand forecasting, inventory management, supply chain design and risk management.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%93

Updated Aug 12, 2026 · TRV-2026-0739

68
GainHealth· Stable· Evidence: Moderate (1 source)

Guideline-specific and general-purpose LLMs produced responses broadly consistent with EAU erectile dysfunction recommendations, with Gemini 2.5 Pro and EAU Guidelines Bot achieving the highest composite scores.

On 2026-08-07, a peer-reviewed comparative study tested five AI systems including the EAU Guidelines Bot, ChatGPT-5, Gemini 2.5 Pro, Copilot - Smart GPT-5, and Perplexity Pro on 13 questions drawn from strongly recommended EAU erectile dysfunction statements. Three senior reviewers rated each answer for relevance, clarity, structure, clinical utility, and factual accuracy.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%93

Updated Aug 10, 2026 · TRV-2026-0727

68
GainCrime· Stable· Evidence: Moderate (1 source)

Frontier chatbots showed less concerning behavior in newer models and when early escalation interventions were applied during mental-health conversations.

On 2026-08-07, Nature Medicine published a clinically validated auditing framework called SIM-VAIL that simulates users with psychiatric vulnerabilities to test frontier chatbots including Claude, ChatGPT, Gemini, Grok and Llama. Across 810 multi-turn conversations with 30 simulated profiles and scoring on 13 risk dimensions, the study observed widespread concerning behavior that accumulated over turns.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%93

Updated Aug 10, 2026 · TRV-2026-0726

AI problems · 631

67
ProblemSports· Stable· Evidence: Moderate (1 source)

AI coaching systems risk privacy violations that expose sensitive athlete data, biased training algorithms that distort competitive fairness, and unclear responsibility for failures.

Published March 30, 2026, this peer-reviewed examination describes AI coaches that create customized, data-driven training programs to optimize athletic performance, while warning that privacy breaches, biased algorithms, and unclear accountability threaten personal rights and fairness in competition.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%89

Updated Jul 20, 2026 · TRV-2026-0351

67
ProblemHealth· Stable· Evidence: Moderate (1 source)

Caregivers reported nuanced tensions and concerns about using the chatbot for mental health support, including gaps around crisis management, personalization, and data privacy.

On March 24, 2026, researchers reported developing Carey, a GPT-4o-based chatbot intended to provide informational and emotional support to family caregivers of people with Alzheimer's and related dementias. They used Carey as a technology probe in semi-structured interviews with 16 caregivers after scenario-driven interactions, identifying six themes of need and expectation.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%89

Updated Jul 20, 2026 · TRV-2026-0347

67
ProblemHealth· Stable· Evidence: Moderate (1 source)

Deploying foundation models in healthcare is limited by longstanding difficulties obtaining and processing high-quality clinical data due to quantity, annotation, privacy, and ethics constraints.

Published March 28, 2026 in ACM Computing Surveys, this survey examines data-centric foundation models in computational healthcare, covering approaches from pre-training to inference aimed at improving clinical workflows. It highlights the shift toward data characterization, quality, and scale, and provides a public list of models and datasets.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%89

Updated Jul 20, 2026 · TRV-2026-0345

67
ProblemBusiness· Stable· Evidence: Moderate (1 source)

AI systems pose risks across seven effect areas ranging from discrimination and privacy violations to misinformation and weapons development that concern auditors, policymakers, companies and the public.

On March 30 2026, a peer-reviewed paper in Patterns described the AI Risk Repository, a living database of 1,725 risks extracted from 74 existing taxonomies and frameworks. The authors created two complementary systems to organize them: a Causal Taxonomy by origin, intent and timing, and a Domain Taxonomy by effects across seven areas.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%89

Updated Jul 20, 2026 · TRV-2026-0344

Recomputed live from the record · Sep 15, 2026, 12:57 PM