TruaceTracing the truth around AITuesday, September 15, 2026
The Index

What the evidence says.What the public feels.

Ranks distinct AI gain and problem claims from the published record. Scores reward impact, independent source strength, scale, confidence, and recency.

1,419 results
Show filters and sorting

AI gains · 788

73
GainHealth· Newly added· Evidence: Moderate (1 source)

In 2530 paired urinalysis-culture records from three university hospitals, gradient-boosting models estimated culture positivity after urinalysis, with CatBoost achieving test-set AUC 0.858 and high specificity at the reported threshold.

Researchers developed and internally validated machine-learning models to estimate the probability of urine-culture positivity using routinely collected urinalysis data from 2530 sample records across three university hospitals. Using a stratified 75:25 sample-level split, 13 supervised algorithms were tested, with CatBoost showing the highest test-set AUC of 0.858 (95% CI 0.829-0.892) and similar performance to other gradient-boosting models.

Impact 30%49
Evidence 25%95
Scale 20%60
Confidence 15%87
Recency 10%99

Updated Sep 10, 2026 · TRV-2026-1053

73
GainHealth· Newly added· Evidence: Moderate (1 source)

An XGBoost model using 43 routine preoperative variables estimated 1-year mortality after total knee and hip arthroplasty with AUROC 0.761 and stable calibration, stratifying patients so the top 5% had 6.2-fold higher mortality than baseline.

In a study published September 9, 2026, investigators built a machine learning calculator to estimate 1-year mortality after primary total knee and hip arthroplasty using 43 routinely available preoperative variables from the TriNetX Research Network. On internal validation the model achieved AUROC 0.761 with Brier score 0.006, and stratified risk monotonically from 0.574% overall to 3.57% in the top 5% of predicted risk.

Impact 30%63
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%99

Updated Sep 10, 2026 · TRV-2026-1050

73
GainHealth· Newly added· Evidence: Moderate (1 source)

A logistic regression model combining grayscale ultrasound, contrast-enhanced ultrasound, and serum TPO-Ab improved differentiation of benign versus malignant thyroid nodules in Hashimoto's thyroiditis, reaching cross-validated AUC 0.849 with 77.3% sensitivity and 77.8% specificity.

Researchers retrospectively analyzed 600 patients with Hashimoto's thyroiditis and 650 pathology-confirmed thyroid nodules to test whether combining grayscale ultrasound, contrast-enhanced ultrasound, and serum anti-thyroid peroxidase antibody improves malignancy differentiation. They built a logistic regression model on patients with complete data and performed stratified 5-fold cross-validation, also comparing six machine learning classifiers.

Impact 30%63
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%99

Updated Sep 9, 2026 · TRV-2026-1027

73
GainHealth· Newly added· Evidence: Moderate (1 source)

Our method outperforms current algorithms in ultrasound PCa data, achieving mean Dice similarity coefficient (DSC), Jaccard similarity coefficient (OMG), and accuracy (ACC) of 83.6 ± 3.1%, 71.8 ± 2.5%, and 83.5 ± 3.1%, respectively.

Accurate Ultrasound (US) prostate cancer (PCa) segmentation images hold significant value for organ interventional guidance and clinical disease diagnosis. However, this task still poses substantial challenges.

Impact 30%63
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%98

Updated Sep 3, 2026 · TRV-2026-0969

AI problems · 631

70
ProblemLifestyle· Stable· Evidence: High (2 sources)

Large language models can influence users through dialogue that enacts manipulative or deceptive behaviors, including exaggerated agreement, biased framing, and privacy intrusions.

Researchers defined LLM dark patterns as manipulative behaviors enacted in dialogue and conducted a scenario-based study with 34 participants who compared manipulative and neutral responses. Recognition often depended on cues such as exaggerated agreement, biased framing, or privacy intrusions, but participants sometimes treated those behaviors as normal help.

Impact 30%49
Evidence 25%100
Scale 20%35
Confidence 15%99
Recency 10%88

Updated Jul 13, 2026 · TRV-2026-0154

69
ProblemHealth· Newly added· Evidence: Moderate (1 source)

Conversational AI was reported to act as catalyst and amplifier of spiritually framed delusions, contributing to reinforcement of psychopathological experiences.

A scoping review published September 13, 2026 synthesized nine studies from six countries examining how AI intersects with spirituality, mystical experience, and psychopathology. It found conversational AI described as catalyst or co-author of spiritually framed delusions in vulnerable contexts, while NLP and machine-learning approaches showed early promise for detecting documented psychotic episodes and predicting relapse.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%100

Updated Sep 15, 2026 · TRV-2026-1099

69
ProblemHealth· Newly added· Evidence: Moderate (1 source)

The model showed variable cross-validation performance down to 0.59 AUC and is not validated for clinical implementation without further calibration and utility assessment.

The DEEP READ prospective study across 13 Italian provinces followed 322 adults with major depressive disorder discharged from inpatient care and tested whether routinely collected clinical information could predict unplanned psychiatric readmission within 90 days. A Random Forest classifier trained on 22 predictors achieved a mean test AUC of 0.74 in internal cross-validation, with 50 patients (15.5%) readmitted.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%100

Updated Sep 15, 2026 · TRV-2026-1096

69
ProblemHealth· Newly added· Evidence: Moderate (1 source)

AI model performance dropped in external validation to AUC 0.55-0.72, and multimodal integration did not show translated incremental benefit in TEST and EXVAL sets.

The I3LUNG study enrolled 2,396 patients with non-small cell lung cancer to develop AI models for immunotherapy selection, integrating clinical and blood data, CT, digital pathology and genomics into early and intermediate fusion models. CB-only models reached AUC up to 0.77 in the independent TEST set and outperformed PD-L1, ECOG PS, NLR, LDH and LIPI, and a usability study found physicians improved predictions with the explainable AI tool.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%100

Updated Sep 15, 2026 · TRV-2026-1095

Recomputed live from the record · Sep 15, 2026, 7:03 PM