Turkish Journal of Geriatrics
2026 , Vol 29, Issue 2
MOST FREQUENTLY ASKED QUESTIONS BY OLDER ADULTS IN GERIATRIC REHABILITATION: EVALUATING LARGE LANGUAGE MODELS AS A SOURCE OF INFORMATION
1Tokat Gaziosmanpasa University, Department of Physical Therapy and Rehabilitation, Tokat, Türkiye2Hatay Mustafa Kemal University, Department of Physical Therapy and Rehabilitation, Hatay, Türkiye
3Bursa Uludag University, Department of Physical Medicine and Rehabilitation, Faculty of Medicine, Bursa, Türkiye
4Bursa Technical University, Graduate School of Education, Bursa, Türkiye DOI : 10.29400/tjgeri.2026.491 Introduction: Older adults in geriatric rehabilitation are increasingly turning to internet-based resources and large language models for health- related information outside of clinical follow-up. The aim of this study was to compare the reliability, clinical accuracy, quality, usefulness, and readability of ChatGPT-5.2, Gemini 3, and DeepSeek V3.2 responses to patient questions regarding geriatric rehabilitation. Materials and Method: In this cross-sectional comparative content analysis, 24 predefined questions on geriatric rehabilitation were developed from YouTube comments, relevant literature, and clinical experience. Each question was submitted to three large language models under standardized conditions. Anonymized responses were independently evaluated by two experienced physiotherapists for reliability, clinical accuracy, quality, and usefulness, with disagreements resolved by consensus with a specialist physician. Readability was assessed using the Flesch Reading Ease scores. Results: No statistically significant difference was found among the three large language models in terms of reliability scores (p = 0.097). However, significant differences were observed for clinical accuracy, quality, usefulness, readability, and text characteristics (all p = 0.001). ChatGPT-5.2 and DeepSeek V3.2 showed the highest clinical accuracy scores, while ChatGPT-5.2 was superior in terms of quality and usefulness. For readability, ChatGPT-5.2 and Gemini 3 outperformed DeepSeek V3.2. Conclusion: Although all three large language models generally produced reliable content, ChatGPT-5.2 and DeepSeek V3.2 showed stronger clinical accuracy performance. Nevertheless, because of the risk of incorrect information being generated, the use of large language models by the older population should preferably be done under expert supervision. Keywords : Geriatrics; Rehabilitation; Artificial Intelligence; Patient Education as Topic; Health Literacy