TruaceTracing the truth around AITuesday, September 15, 2026
The Index

What the evidence says.What the public feels.

Ranks distinct AI gain and problem claims from the published record. Scores reward impact, independent source strength, scale, confidence, and recency.

1,419 results
Show filters and sorting

AI gains · 788

72
GainHealth· Stable· Evidence: Moderate (1 source)

Large language models directed patients with musculoskeletal complaints to currently practicing, specialty-appropriate providers in the requested city, with ChatGPT achieving 100% appropriateness in the tested queries.

Researchers prompted ChatGPT, DeepSeek, and Gemini with standardized musculoskeletal complaints for Lynchburg, VA and Trumbull, CT, and judged whether recommended physicians were currently practicing locally in the relevant specialty and whether phone numbers were correct. By the August 13, 2026 publication date, ChatGPT was appropriate in all 17 recommendations, while Gemini and DeepSeek were appropriate in 43% and 40% of recommendations respectively.

Impact 30%63
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%94

Updated Aug 14, 2026 · TRV-2026-0752

72
GainHealth· Stable· Evidence: Moderate (1 source)

YOLOv8x-pose and YOLOv8x-seg models automated landmark detection and apical segmentation for the Cameriere European method in children aged 5-13, achieving high detection accuracy and low measurement error in a first-stage validation.

This first-stage retrospective validation evaluated YOLOv8-based models to automate the anatomical inputs for the Cameriere European dental age estimation method using 4,050 panoramic radiographs of children aged 5-13. A YOLOv8x-pose model detected open-apex landmarks and a YOLOv8x-seg model segmented closed apices, with performance compared to manual reference annotations from CranioCatch software.

Impact 30%63
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%93

Updated Aug 9, 2026 · TRV-2026-0715

72
GainHealth· Stable· Evidence: Moderate (1 source)

ChatGPT-5.0 achieved 75.77% accuracy and the highest sensitivity at 76.68% for orthodontic extraction decisions, performing comparably to XGBoost and significantly better than random forest, SVM, logistic regression and MLP.

A comparative study published August 7, 2026 evaluated ChatGPT-5.0 against five supervised machine learning algorithms for orthodontic extraction decisions. Using 520 cases (42.88% extraction, 57.12% non-extraction) and 23 clinical, cephalometric and photographic variables, with expert consensus as reference, ChatGPT-5.0 achieved 75.77% accuracy and 76.68% sensitivity under 5-fold cross-validation, compared to 78.08% accuracy for XGBoost.

Impact 30%63
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%93

Updated Aug 8, 2026 · TRV-2026-0689

72
GainHealth· Stable· Evidence: Moderate (1 source)

K-means clustering of eight biopsychosocial variables identified three distinct AUD profiles that predicted 3-month abstinence, with Late-Onset achieving 65.8% abstinence, supporting stratified front-loaded intervention.

In a retrospective study of 102 patients at a tertiary care center in India, researchers used k-means clustering on eight biopsychosocial baseline variables to derive three AUD profiles. By the August 2026 publication date, they reported Late-Onset, High-Functioning, and Severe groups with differing 3-month abstinence rates corroborated by GGT levels and bootstrap-assessed cluster stability.

Impact 30%63
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%93

Updated Aug 8, 2026 · TRV-2026-0687

AI problems · 631

68
ProblemHealth· Newly added· Evidence: Moderate (1 source)

Even when incorrect, ChatGPT and Claude produced thorough justifications, creating risk of convincing but wrong guideline-based information, while Gemini showed formatting deviations.

A peer-reviewed study published September 7, 2026 compared three large language models to 10 international hip preservation experts on a 21-item questionnaire derived from consensus guidelines for femoroacetabular impingement, dysplasia and microinstability. Gemini scored 100%, ChatGPT 98.4% and Claude 96.8% versus 90.5% for experts, with models also showing higher intra-item agreement.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%99

Updated Sep 8, 2026 · TRV-2026-1018

68
ProblemEducation· Newly added· Evidence: Moderate (1 source)

Educators and students report concern about misuse of GenAI in education and a lack of classroom and institutional preparedness to manage it in the writing process.

In a November 2023 peer-reviewed study, researchers surveyed 68 educators and 158 university students about when generative AI should be used in academic writing. Participants reviewed ChatGPT prompts and outputs for brainstorming, outlining, writing, revising, feedback, and evaluating, and rated acceptability for student versus teacher use.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%98

Updated Sep 7, 2026 · TRV-2026-1012

68
ProblemHealth· Newly added· Evidence: Moderate (1 source)

AI-driven personalized nutrition is limited by algorithmic bias, poor generalizability, and data privacy risks that prevent fair and reliable clinical application.

By October 2025, peer-reviewed literature described AI as a key enabler of personalized nutrition, using machine learning to analyze multiomics datasets and guide microbiome-based dietary interventions for obesity, diabetes, cardiovascular and gastrointestinal disorders with support from digital twins and health knowledge graphs.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%98

Updated Sep 7, 2026 · TRV-2026-1011

68
ProblemHealth· Newly added· Evidence: Moderate (1 source)

Model generalizability for non-displaced Type I fractures is limited by small sample size, and pooled evidence remains preliminary with substantial heterogeneity across studies.

Researchers developed a YOLOv11 Nano model to detect pediatric supracondylar fractures and classify Gartland subtypes I-III on 1082 elbow radiographs from 2004-2018, testing three patient-level validation schemes and adding bone segmentation and explainable AI visualizations. They also conducted a PRISMA-DTA meta-analysis of four studies totaling 2232 images comparing CNN and radiomics approaches.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%98

Updated Sep 7, 2026 · TRV-2026-1008

Recomputed live from the record · Sep 15, 2026, 9:59 PM