TruaceTracing the truth around AITuesday, September 15, 2026
The Index

What the evidence says.What the public feels.

Ranks distinct AI gain and problem claims from the published record. Scores reward impact, independent source strength, scale, confidence, and recency.

1,419 results
Show filters and sorting

AI gains · 788

74
GainHealth· Newly added· Evidence: Moderate (1 source)

FedMediFormer-XAI combined federated learning, multimodal transformers, diffusion augmentation and GNN drug recommendation to achieve 94.2% accuracy for diabetes prediction and NDCG 0.91 for personalized drug recommendation while enabling privacy-preserving training without sharing raw patient data.

Researchers described FedMediFormer-XAI, a framework that combines federated learning, multimodal transformers, diffusion-based data augmentation, graph neural networks for drug recommendation, and explainable AI to address fragmented diabetes data, privacy concerns, and lack of personalized guidance. It was tested on diverse inputs including clinical records, glucose monitoring, retinal fundus images, wearable sensors, and pharmacological information.

Impact 30%69
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%99

Updated Sep 10, 2026 · TRV-2026-1052

74
GainHealth· Newly added· Evidence: Moderate (1 source)

In a 20-case test of congenital anomalies and tumors, ChatGPT, Gemini and Copilot achieved 80-95% diagnostic accuracy with detailed anatomical descriptions, suggesting potential as supplementary radiological diagnostic support.

On September 9, 2026, a peer-reviewed comparative study reported testing ChatGPT, Gemini, and Microsoft Copilot on 20 radiological cases split between congenital anomalies and tumors. Each system received the same questions and images and was scored for diagnostic accuracy and explanatory completeness, with ChatGPT scoring 19 correct, Gemini 17, and Copilot 16.

Impact 30%69
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%99

Updated Sep 10, 2026 · TRV-2026-1047

74
GainHealth· Newly added· Evidence: Moderate (1 source)

An ensemble deep learning model combining AP and lateral skull X-rays detected skull fractures in neonates and infants with 91.6% accuracy and 0.938 AUC on external validation, improving diagnostic accuracy while reducing need for CT.

By September 7, 2026, researchers reported developing an AI-assisted ensemble model to detect skull fractures in neonates and infants from plain radiographs. Using a retrospective set of 1,184 patients from 2010-2021, they preprocessed images with CLAHE and trained three CNNs, with DenseNet-121 performing best, then combined AP and lateral views into an ensemble that achieved 0.938 AUC and 91.6% accuracy on an external set of 460 images.

Impact 30%69
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%99

Updated Sep 9, 2026 · TRV-2026-1031

74
GainHealth· Newly added· Evidence: Moderate (1 source)

Integrating surface-enhanced Raman spectroscopy with a support vector machine model enabled rapid species-level identification of seven clinically common Nocardia spp. at 99.47% accuracy to guide clinical treatment.

On 2026-09-08, a peer-reviewed study reported an intelligent analytical model combining surface-enhanced Raman spectroscopy with machine learning to identify seven clinically common Nocardia species from cultured clinical isolates. Using 46 strains and 64 spectra per strain, the team compared nine models and found the support vector machine achieved 99.47% accuracy.

Impact 30%69
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%99

Updated Sep 9, 2026 · TRV-2026-1029

AI problems · 631

72
ProblemHealth· Stable· Evidence: Moderate (1 source)

Some ambiguous cases remained unresolved when investigator replies lacked actionable information, and cases with no reply after 48 hours still required deferral to offline manual review.

Researchers developed TrialTriage, a semiautonomous prescreening workflow on the n8n platform that uses large language model extraction from clinical narratives and investigator email replies plus a 7-criterion deterministic rule engine to classify phase I oncology trial eligibility, automatically emailing investigators when information is missing and reclassifying after reply capture.

Impact 30%63
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%92

Updated Aug 6, 2026 · TRV-2026-0666

72
ProblemHealth· Stable· Evidence: Moderate (1 source)

AI-driven diabetic retinopathy screening using ultra-widefield fundus images showed limited specificity of 72.5% in meta-analysis.

A systematic review and meta-analysis to February 9, 2025, evaluated artificial intelligence for diabetic retinopathy assessment using ultra-widefield color fundus images, which capture a larger retinal area without pupil dilation. Of 527 records, 17 studies were reviewed and four were meta-analyzed, all using Optos software, to estimate sensitivity and specificity for AI-driven screening.

Impact 30%63
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%90

Updated Jul 25, 2026 · TRV-2026-0564

72
ProblemHealth· Stable· Evidence: Moderate (1 source)

LLM summarisation of consultations showed a 1.47% hallucination rate and 3.45% omission rate, creating fidelity gaps that could compromise patient safety.

On 2025-05-13, a peer-reviewed framework was described for evaluating LLMs that automate summarising consultations into clinical notes. It combines an error taxonomy, iterative experimental comparisons, a clinical safety harm assessment, and the CREOLA interface, tested across 18 configurations with 12,999 clinician-annotated sentences.

Impact 30%63
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%90

Updated Jul 24, 2026 · TRV-2026-0522

72
ProblemLifestyle· Stable· Evidence: Moderate (1 source)

Deployment of ML for food quality control is constrained by data scarcity, domain coverage biases, and challenges integrating with legacy systems while meeting regulatory compliance and cost-benefit requirements.

On 2025-10-04, a peer-reviewed review in Foods synthesized 25 studies selected from 124 Scopus records from 2005-2025 to map machine learning use for quality control in food production. It organized findings into six domains covering quality applications, defect detection and visual inspection, ingredient optimization, packaging sensors and predictive QC, supply chain traceability, and Industry 4.0 models.

Impact 30%49
Evidence 25%95
Scale 20%60
Confidence 15%87
Recency 10%89

Updated Jul 22, 2026 · TRV-2026-0481

Recomputed live from the record · Sep 15, 2026, 3:13 PM