TruaceTracing the truth around AIMonday, September 14, 2026

All traces

Temporal and cross-site validation of an AI system for self-harm detection
HealthContested · G 69 / P 71

AI system for self-harm detection in emergency department triage notes

Source article: Temporal and cross-site validation of an AI system for self-harm detection

Problem

When applied to a regional hospital 150 km outside Melbourne, the same AI system's ability to distinguish self-harm cases declined to PR AUC 0.78, with instability linked to linguistic domain shift and different self-harm presentations.

PLOS Digital Health
Gain

An AI system combining text normalisation with 1931 features maintained stable self-harm detection in prospective validation at its development metropolitan hospital, achieving PR AUC 0.84 over 329,655 triage notes in the following four years.

PLOS Digital Health
Explainable machine learning for breast cancer prediction in resource-constrained settings: A multi-algorithmic framework integrating shap-based transparency with clinical decision support
HealthContested · G 71 / P 71

machine learning for breast cancer diagnosis in resource-constrained settings

Source article: Explainable machine learning for breast cancer prediction in resource-constrained settings: A multi-algorithmic framework integrating shap-based transparency with clinical decision support

Problem

Algorithmic opacity and lack of interpretability frameworks tailored to resource-constrained environments have impeded clinical adoption of ML for breast cancer, contributing to diagnostic delays in settings with limited pathology capacity.

PLOS Digital Health
Gain

Explainable ML models achieved near-perfect discrimination for breast cancer diagnosis on cytology data, with top models reaching 0.996 AUC and 98.25% accuracy, supporting use in resource-constrained diagnostic workflows.

PLOS Digital Health
Comparative quality, accuracy, and readability of large language model responses to patient questions about robotic-assisted total knee arthroplasty
HealthContested · G 70 / P 72

LLM-generated patient information about robotic-assisted total knee arthroplasty

Source article: Comparative quality, accuracy, and readability of large language model responses to patient questions about robotic-assisted total knee arthroplasty

Problem

LLM-generated answers to patient questions about robotic-assisted total knee arthroplasty remained above recommended patient-education reading levels and should be regarded as supplementary rather than standalone sources of information.

The Knee
Gain

Large language models including ChatGPT-o3, ChatGPT-5.2, Gemini 3 and DeepSeek provided generally acceptable clinical accuracy when answering 30 frequently asked patient questions about robotic-assisted total knee arthroplasty.

The Knee
Systematic Bias in Comparative Evaluations of Machine Learning Versus Logistic Regression for Clinical Prediction Models: A Meta-Research Analysis Using Trauma Mortality as an Empirical Case
HealthContested · G 69 / P 68

comparative discrimination of machine learning versus logistic regression for trauma mortality prediction on same datasets

Source article: Systematic Bias in Comparative Evaluations of Machine Learning Versus Logistic Regression for Clinical Prediction Models: A Meta-Research Analysis Using Trauma Mortality as an Empirical Case

Problem

Apparent superiority of ML over logistic regression for trauma mortality prediction may be inflated by convergent practices including comparing best-of-several ML models to a single LR comparator, reliance on internal validation, selective reporting, and AUC-only synthesis.

Journal of Clinical Epidemiology
Gain

Across 17 studies totaling 243,324 trauma patients, the best-performing ML model showed a small pooled AUC advantage over logistic regression for mortality prediction.

Journal of Clinical Epidemiology
Predicting Conversion from Mild Cognitive Impairment to Alzheimer's Disease: A Systematic Review of Deep Learning Models for Early-Stage Disease Classification
HealthContested · G 74 / P 72

deep learning models predicting conversion from mild cognitive impairment to Alzheimer's disease

Source article: Predicting Conversion from Mild Cognitive Impairment to Alzheimer's Disease: A Systematic Review of Deep Learning Models for Early-Stage Disease Classification

Problem

Deep learning models for MCI-to-AD conversion face substantial barriers to routine clinical use due to heavy reliance on ADNI, lack of diverse multicenter data, overfitting, and poor interpretability.

Ageing Research Reviews
Gain

Deep learning models showed promising and often strong performance for predicting conversion from mild cognitive impairment to Alzheimer's disease, supporting early diagnosis and timely therapeutic intervention.

Ageing Research Reviews
Gavin Newsom imposes strict new rules on AI, social media and chatbots for children
PolicyContested · G 60 / P 59

California AB 1709 and companion chatbot regulations restricting social media features for users under 16

Source article: Gavin Newsom imposes strict new rules on AI, social media and chatbots for children

Problem

Critics argue AB 1709 functions as a ban for under-16s that cuts young people off from digital lifelines and speech without making them safer or healthier.

The Guardian
Gain

California's new package creates the strongest U.S. companion chatbot rules and bans addictive features like infinite scroll for under-16s to protect young users from harm.

The Guardian
The impact of digital technology, social media, and artificial intelligence on cognitive functions: a review
HealthContested · G 67 / P 64

effects of digital devices, social media and AI tools on cognitive functions including attention and memory

Source article: The impact of digital technology, social media, and artificial intelligence on cognitive functions: a review

Problem

Digital devices, social media and AI tools influence brain function and cognitive abilities, with potential negative impacts on attention, memory and related functions.

Frontiers in Cognition
Gain

Digital devices, social media and AI tools have brought convenience and connectivity and can have positive impacts on cognitive functions including attention and memory.

Frontiers in Cognition
Machine-learning prediction of urine-culture positivity in a multicentre test-ordered cohort: Model development and internal validation
HealthContested · G 68 / P 70

machine-learning prediction of urine-culture positivity from routine urinalysis data in patients with paired orders

Source article: Machine-learning prediction of urine-culture positivity in a multicentre test-ordered cohort: Model development and internal validation

Problem

Models were validated only at sample level without patient or centre grouping, and exploratory risk strata were not evaluated for clinical utility or safety, so they do not establish symptomatic UTI or safe antibiotic decisions.

BJUI Compass
Gain

In 2530 paired urinalysis-culture records from three university hospitals, gradient-boosting models estimated culture positivity after urinalysis, with CatBoost achieving test-set AUC 0.858 and high specificity at the reported threshold.

BJUI Compass
Machine Learning for Mortality Prediction in Infective Endocarditis: A Systematic Review and Meta-Analysis
HealthContested · G 69 / P 73

supervised ML models predicting all-cause mortality in adult infective endocarditis patients

Source article: Machine Learning for Mortality Prediction in Infective Endocarditis: A Systematic Review and Meta-Analysis

Problem

Half of included studies had identified risk of bias and clinical adoption remains limited, requiring multicenter prospective validation and interpretable frameworks before bedside use.

Cardiology in Review
Gain

Supervised ML models, especially ensemble methods, predicted all-cause mortality in adult infective endocarditis with pooled AUC 0.85 for both in-hospital/early and 6-month mortality, outperforming conventional scores.

Cardiology in Review
Artificial intelligence for lung disease quantification in systemic sclerosis-associated interstitial lung disease and other connective tissue disease-associated interstitial lung disease
HealthNegative state · G 67 / P 76

AI-based HRCT quantification for systemic sclerosis-associated interstitial lung disease risk stratification and clinical decision support

Source article: Artificial intelligence for lung disease quantification in systemic sclerosis-associated interstitial lung disease and other connective tissue disease-associated interstitial lung disease

Problem

Visual HRCT scoring remains reader-dependent and AI outputs lack prospective multicenter validation and protocol harmonization needed to serve as treatment-triggering biomarkers.

Current Opinion in Rheumatology
Gain

In systemic sclerosis-associated ILD, AI-based HRCT quantification stratifies FVC decline and long-term survival and correlates with lung function measures to predict mortality.

Current Opinion in Rheumatology
How Well Do AI Chatbots Understand Abnormal Anatomy: A Comparative Study Using Congenital Anomalies and Tumor Cases
HealthContested · G 70 / P 70

AI chatbot interpretation of radiological congenital anomaly and tumor cases

Source article: How Well Do AI Chatbots Understand Abnormal Anatomy: A Comparative Study Using Congenital Anomalies and Tumor Cases

Problem

Chatbots sometimes confused similar congenital anomalies and provided less detailed anatomical descriptions in complex tumor cases, requiring caution and verification by qualified professionals before clinical use.

Clinical Anatomy
Gain

In a 20-case test of congenital anomalies and tumors, ChatGPT, Gemini and Copilot achieved 80-95% diagnostic accuracy with detailed anatomical descriptions, suggesting potential as supplementary radiological diagnostic support.

Clinical Anatomy
Stakeholder perspectives on artificial intelligence in schizophrenia care
HealthContested · G 70 / P 69

use of a hypothetical AI companion tool for people with schizophrenia spectrum disorders

Source article: Stakeholder perspectives on artificial intelligence in schizophrenia care

Problem

Participants identified trust as the central barrier to using an AI companion, driven by privacy concerns and vulnerabilities specific to schizophrenia.

Psychological Medicine
Gain

Participants with schizophrenia recognized an AI companion as a potentially accessible source of support between clinical visits.

Psychological Medicine
Effect of Large Language Model-Powered Virtual Standardized Patients on History-Taking Among Undergraduate Medical Students: Propensity-Matched Cohort Study
HealthContested · G 71 / P 71

LLM-VSP self-practice effect on undergraduate medical students' medical history-taking performance

Source article: Effect of Large Language Model-Powered Virtual Standardized Patients on History-Taking Among Undergraduate Medical Students: Propensity-Matched Cohort Study

Problem

Students with medium and low baseline history-taking proficiency showed relatively limited score improvements from LLM-VSP self-practice, with practice frequency alone not independently predicting final performance.

JMIR Medical Education
Gain

Undergraduate medical students who used LLM-powered virtual standardized patients as extracurricular self-practice achieved higher end-of-term history-taking performance at an OSCE with real standardized patients compared to routine instruction.

JMIR Medical Education
Beyond the Algorithm: A Stewardship Framework for the Hand Surgeon Adopting Artificial Intelligence
HealthContested · G 70 / P 73

AI tools for hand surgery imaging, outcome prediction, and communication affecting hand surgery patients and clinical outcomes

Source article: Beyond the Algorithm: A Stewardship Framework for the Hand Surgeon Adopting Artificial Intelligence

Problem

Most hand surgery AI tools are deployed in unaudited workflows and rarely remeasured after release after testing only on training-like data, leaving the hand surgeon accountable for patient outcomes shaped by opaque models.

The Journal of Hand Surgery
Gain

AI tools are entering hand surgery practice to read scaphoid and distal radius radiographs and to predict outcomes after carpal tunnel release.

The Journal of Hand Surgery
A Quality Assessment Rubric for Artificial Intelligence-Generated Patient-Friendly Radiology Reports
HealthContested · G 69 / P 73

safety and quality of AI-generated patient-friendly radiology reports for patient distribution

Source article: A Quality Assessment Rubric for Artificial Intelligence-Generated Patient-Friendly Radiology Reports

Problem

AI tools translating radiology reports into plain language can produce translation errors that compromise comprehension and safety, causing reports to be graded unsafe and warrant withholding from patients.

American Journal of Roentgenology
Gain

A five-attribute rubric for AI-generated patient-friendly radiology reports showed almost-perfect agreement between lay and radiologist team members and may provide a standardized safeguard before patient distribution.

American Journal of Roentgenology