TruaceTracing the truth around AITuesday, September 15, 2026
TRV-2026-1095Version 1 · Certified

Written 2026-09-15 06:55:45 UTC · current record

Reason for this version

Certified into the record

Canonical text (the exact bytes fingerprinted)

TRUVACE RECORD VERSION
record: TRV-2026-1095
version: 1
kind: certified
reason: Certified into the record
timestamp: 2026-09-15T06:55:45.305811Z
status: published
lens: trace
sector: health
headline: Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC
dek: Despite a decade in, immunotherapy (IO) treatment selection in non-small cell lung cancer (NSCLC) remains largely guided by subgroup analyses and imperfect programmed death ligand 1 (PD-L1) and clinical scores. To our knowledge, I 3 LUNG ( NCT05537922 ) is currently the largest international, real-world, multimodal, artificial intelligence (AI)-based study, enrolling 2,396 patients. We integrated real-world clinical and blood (CB) data, computed tomography (CT) images, digital pathology (DP), and genomics into m…
gain_title: In NSCLC immunotherapy selection, CB-only AI models outperformed PD-L1 and clinical scores in the independent test set, and both expert and nonexpert physicians improved their predictions when using the explainable AI decision support tool.
problem_title: AI model performance dropped in external validation to AUC 0.55-0.72, and multimodal integration did not show translated incremental benefit in TEST and EXVAL sets.
trace_subject: AI-based prediction of immunotherapy outcomes in non-small cell lung cancer using clinical, blood, imaging, pathology and genomic data
gain_reading: In NSCLC immunotherapy selection, CB-only AI models outperformed PD-L1 and clinical scores in the independent test set, and both expert and nonexpert physicians improved their predictions when using the explainable AI decision support tool.
gain_evidence: AI models significantly surpassed PD-L1, Eastern Cooperative Oncology Group performance status (ECOG PS), neutrophil-to-lymphocyte ratio (NLR), lactate dehydrogenase (LDH) and Lung Immune Prognostic Index (LIPI) score in the independent TEST set | lung expert and nonexpert physicians improved their prediction with the explainable AI (XAI) ML CB-only based tool
problem_reading: AI model performance dropped in external validation to AUC 0.55-0.72, and multimodal integration did not show translated incremental benefit in TEST and EXVAL sets.
problem_evidence: Performance drop in external validation (EXVAL) likely reflects population differences (AUC range: 0.55-0.72) | Although multimodal integration with MLEF (CB+CT+DP) was associated with higher performance, its incremental benefit remains uncertain, not translated in TEST and EXVAL
quick_read: The I3LUNG study enrolled 2,396 patients with non-small cell lung cancer to develop AI models for immunotherapy selection, integrating clinical and blood data, CT, digital pathology and genomics into early and intermediate fusion models. CB-only models reached AUC up to 0.77 in the independent TEST set and outperformed PD-L1, ECOG PS, NLR, LDH and LIPI, and a usability study found physicians improved predictions with the explainable AI tool.

The findings matter because immunotherapy selection still relies on imperfect biomarkers, and an explainable tool that helps both experts and nonexperts could change practice if validated. Uncertainty remains due to lower AUC of 0.55-0.72 in external validation and lack of translated benefit from multimodal integration, with prospective validation in over 2,000 patients still ongoing at publication.
limitation: External generalizability is limited by performance drop in external validation and uncertain incremental benefit of multimodal fusion, which did not translate in TEST and EXVAL.
tag: Dual reading
key_points: I3LUNG (NCT05537922) enrolled 2,396 patients as largest international real-world multimodal AI study in NSCLC. | CB-only ML and DL models reached AUC up to 0.77 in TEST set. | Clinical usability study tested XAI ML CB-only tool with lung expert and nonexpert physicians. | Prospective validation of decision support system in more than 2,000 patients was ongoing as of publication date.
rundown: The study integrated real-world clinical and blood data, CT images, digital pathology, and genomics into MLEF and DLIF models, with CB-only models achieving AUC up to 0.77 in TEST.

In the clinical usability study, both expert and nonexpert physicians improved prediction using the XAI ML CB-only tool, while multimodal CB+CT+DP fusion showed higher performance in development but uncertain benefit in independent evaluation.
sources:
- peer_reviewed | Nature Medicine | https://doi.org/10.1038/s41591-026-04488-2 | 2026-09-13
prev: 0000000000000000000000000000000000000000000000000000000000000000
sha256
439de0637552e79cdd80e8247a34ebdb0543fcbc1248c0863a39e691521fe0bf
previous
0000000000000000000000000000000000000000000000000000000000000000
Verify this record
How to verify without trusting this page

Fetch the canonical text of any version from /api/record/TRV-2026-1095 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.