Comparative quality, accuracy, and readability of large language model responses to patient questions about robotic-assisted total knee arthroplasty
Purpose To compare the information quality, accuracy, and readability of patient-directed responses generated by large language models (LLMs), including ChatGPT-o3, ChatGPT-5.2, Gemini 3, and DeepSeek, regarding robotic-assisted total knee arthroplasty (RA-TKA). Methods Thirty frequently asked patient questions were identified using LLM outputs and Google search queries. Responses were evaluated for information quality using the DISCERN and Quality Analysis of Medical Artificial Intelligence (QAMAI) instruments,…
Large language models including ChatGPT-o3, ChatGPT-5.2, Gemini 3 and DeepSeek provided generally acceptable clinical accuracy when answering 30 frequently asked patient questions about robotic-assisted total knee arthroplasty.
LLM-generated answers to patient questions about robotic-assisted total knee arthroplasty remained above recommended patient-education reading levels and should be regarded as supplementary rather than standalone sources of information.
Findings are limited to 30 selected questions and show readability remains above recommended patient-education levels, with no significant difference on QAMAI and authors concluding responses should be supplementary only.
Evidence
- Peer-reviewedThe Knee2026-09-11
How should this claim be treated?
Truvace Impact Record TRV-2026-1069, v1: “Comparative quality, accuracy, and readability of large language model responses to patient questions about robotic-assisted total knee arthroplasty.” Truvace, 2026-09-13. /record/TRV-2026-1069 (accessed at citation time). sha256 9881c86c0a3df0bc…
Calibration history
Every change to this record since certification, in the open. None yet — the reading has held since it entered the record.
Certified into the record
How to verify without trusting this page
Fetch the canonical text of any version from /api/record/TRV-2026-1069 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.
ace