TruaceTracing the truth around AITuesday, September 15, 2026
The Index

What the evidence says.What the public feels.

Ranks distinct AI gain and problem claims from the published record. Scores reward impact, independent source strength, scale, confidence, and recency.

1,419 results
Show filters and sorting

AI gains · 788

74
GainHealth· Stable· Evidence: Moderate (1 source)

TrialScout using a large language model matched registered trials to result publications with 92.5% sensitivity and located publications for 63.6% of a 9,600-trial sample, accelerating locating trial results.

On 2026-08-26, a peer-reviewed study described TrialScout, a program that uses a large language model to link ClinicalTrials.gov registrations to PubMed result publications. Tested against prior human-coded datasets, it achieved 92.5% sensitivity and 81.2% specificity, and when applied to 9,600 sampled completed or terminated trials it identified publications for 6,110 trials.

Impact 30%69
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%96

Updated Aug 28, 2026 · TRV-2026-0918

74
GainHealth· Stable· Evidence: Moderate (1 source)

A three-stage explainable framework that fuses deep representations from multiple pretrained CNNs and classifies them with an RBF-kernel SVM achieved 88% accuracy and 88% F1 for skin lesion classification while adding LIME-based explanations.

Researchers developed EDLF-SLC, an explainable deep learning framework for skin lesion classification that fuses features from multiple pretrained CNNs, classifies them with an RBF-kernel SVM, and explains predictions with LIME. In experiments reported in August 2026, the system reached 88% accuracy and 88% F1, presented as competitive with recent methods.

Impact 30%69
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%96

Updated Aug 27, 2026 · TRV-2026-0905

74
GainHealth· Stable· Evidence: Moderate (1 source)

Dynamic class-specific F1-score weighting with ECF1V increased ensemble classification accuracy to 98.25% on Breast Cancer Wisconsin and 89.47% on UCI Heart Disease datasets, outperforming conventional voting under non-linear and imbalanced conditions.

Researchers developed three adaptive ensemble voting methods that assign weights based on per-class F1-scores from validation instead of overall accuracy. They tested the approaches on Gaussian Mixture, Spiral, and Moon synthetic datasets and on the Breast Cancer Wisconsin and UCI Heart Disease datasets, comparing against majority, weighted, and soft voting.

Impact 30%69
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%96

Updated Aug 26, 2026 · TRV-2026-0895

74
GainCrime· Stable· Evidence: Moderate (1 source)

Supervised logistic regression distinguished dominant versus non-dominant handwriting with 92.98% test accuracy using execution-quality features, providing quantitative benchmarks to support forensic evaluation of suspected off-hand disguise.

Researchers collected paired samples from 94 right-handed participants who each wrote the same standardized text with dominant and non-dominant hands, then scored 13 general and 19 individual characteristics. They found statistically significant differences in 61.5% of general and 57.9% of individual characteristics and trained a supervised logistic regression model to distinguish hand use.

Impact 30%69
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%96

Updated Aug 26, 2026 · TRV-2026-0889

AI problems · 631

72
ProblemClimate· Stable· Evidence: Moderate (1 source)

Expanding AI and digital infrastructure for automation increases demand for computing energy, raising concerns about data center efficiency and carbon footprint under the Green AI vs Red AI divide.

By March 2026, this peer-reviewed overview described how AI-based automation in Industry 4.0 optimized production, logistics, and resource management to reduce waste and energy use, and how Industry 5.0 expanded that with human-machine collaboration, generative AI, digital twins, and decentralized smart grids and microgrids.

Impact 30%49
Evidence 25%95
Scale 20%60
Confidence 15%87
Recency 10%89

Updated Jul 19, 2026 · TRV-2026-0281

72
ProblemHealth· Stable· Evidence: Moderate (1 source)

First-year medical students changed answers to match ChatGPT in 22.3% of cases, with greater reliance on foundational than clinical questions, indicating context-dependent overreliance risk.

In a July 2026 peer-reviewed study, 57 first-year medical students completed 24 paired clinical and foundational questions during a pediatric nephrology and urology case-based session, answering individually, then viewing a ChatGPT-generated answer that was deliberately correct or incorrect, and re-answering.

Impact 30%63
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%88

Updated Jul 17, 2026 · TRV-2026-0235

72
ProblemSports· Stable· Evidence: Moderate (1 source)

Practical utility of machine learning in sport was often limited by issues of data quality, interpretability, and accessibility for end users such as athletes and coaches.

A scoping review published May 25, 2026 examined 270 peer-reviewed studies from 2002 to 2024 on machine learning in sport. It found applications across 12 subject areas, most frequently computer science, biomechanics, and sport psychology, with common uses in action recognition, injury prediction/prevention, and athlete selection/talent identification.

Impact 30%49
Evidence 25%95
Scale 20%60
Confidence 15%87
Recency 10%88

Updated Jul 13, 2026 · TRV-2026-0186

72
ProblemHealth· Stable· Evidence: Moderate (1 source)

In a 960-response vignette test, ChatGPT Health undertriaged 52% of gold-standard emergencies, directing diabetic ketoacidosis and impending respiratory failure to 24-48 h evaluation instead of the emergency department, with failures concentrated at clinical extremes and triage shifting toward less urgent care when by-

In a structured stress test published February 23, 2026, researchers evaluated ChatGPT Health, OpenAI's consumer health tool launched in January 2026, using 60 clinician-authored vignettes across 21 clinical domains under 16 factorial conditions to generate 960 responses, assessing triage recommendations and contextual sensitivity.

Impact 30%63
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%88

Updated Jul 13, 2026 · TRV-2026-0181

Recomputed live from the record · Sep 15, 2026, 4:07 PM