Chinese Large Language Models for Patient-Friendly Knee MRI Summaries in Sports Injury: A Guideline-Anchored Study of Quality and Robustness

Authors

DOI:

https://doi.org/10.65455/2c4jes43

Keywords:

Large Language Models, Knee Magnetic Resonance Imaging, Sports Injury, Plain-Language Reporting, Patient Communication, Readability, Semantic Fidelity, Robustness

Abstract

Reports describing sports-related knee MRI findings are usually written for clinicians and can be difficult for patients to interpret. We compared the quality and consistency of five widely used Chinese large language models (LLMs) when converting such reports into plain language. A guideline-anchored set of 50 simulated knee MRI reports was prepared, and the five models were tested with zero-shot prompting; DeepSeek-V4 was also tested with a few-shot prompt. The six experimental arms produced 300 summaries through the vendors' official web interfaces. Two blinded physicians rated factual accuracy (0-3), information coverage (0-5), and patient-oriented clarity (0-5). We additionally calculated mean readability grade level (mRGL) and BERT-based semantic similarity. Accuracy remained high across arms (2.60-2.83) without a significant pairwise difference after correction. Few-shot DeepSeek-V4 obtained the best coverage score (4.33 +/- 0.61), Doubao-Seed-2.0 generated the easiest text to read (mRGL 7.15), and GLM-5.2 remained closest to the wording of the source reports (similarity 0.72). The models were less alike in difficult combined injuries: Doubao-Seed-2.0 and Qwen3.8-Max showed the clearest decline, while few-shot DeepSeek-V4 had the lowest accuracy CV (12.7%). The findings indicate that Chinese flagship LLMs can produce generally accurate patient-facing knee MRI explanations, but they occupy different positions with respect to coverage, readability, fidelity, and robustness. Safe selection therefore requires more than a single average score.

References

[1]MUSAH L V, KARLSSON J. Anterior cruciate ligament tear. New England Journal of Medicine, 2019, 380(24): 2341-2348. DOI: https://doi.org/10.1056/NEJMcp1805931

[2]KAEDING C C, LEGER-ST-JEAN B, MAGNUSSEN R A. Epidemiology and diagnosis of anterior cruciate ligament injuries. Clinics in Sports Medicine, 2017, 36(1): 1-8. DOI: https://doi.org/10.1016/j.csm.2016.08.001

[3]BROPHY R H, LOWRY K J. Management of anterior cruciate ligament injuries: AAOS clinical practice guideline summary. Journal of the American Academy of Orthopaedic Surgeons, 2023.

[4]American College of Radiology. ACR Appropriateness Criteria: acute trauma to the knee[Z/OL]. https://acsearch.acr.org/docs/69436/Narrative/. Accessed August 19, 2026.

[5]NIELSEN-BOHLMAN L, PANZER A M, KINDIG D A, eds. Health Literacy: A Prescription to End Confusion. Washington, D.C.: National Academies Press, 2004. DOI: https://doi.org/10.17226/10883

[6]PAASCHE-ORLOW M K, WOLF M S. The causal pathways linking health literacy to health outcomes. American Journal of Health Behavior, 2007, 31(suppl 1): S19-S26. DOI: https://doi.org/10.5993/AJHB.31.s1.4

[7]Office of the National Coordinator for Health Information Technology. 21st Century Cures Act: interoperability, information blocking, and the ONC Health IT Certification Program. Federal Register, 2020, 85(85): 25642-25961.

[8]SALLAM M. ChatGPT utility in healthcare education, research, and practice: systematic review on the promising perspectives and valid concerns. Healthcare, 2023, 11(6): 887. DOI: https://doi.org/10.3390/healthcare11060887

[9]JEBLICK K, SCHACHTNER B, DEXL J, et al. ChatGPT makes medicine easy to swallow: an exploratory case study on simplified radiology reports. European Radiology, 2024, 34(5): 2817-2825. DOI: https://doi.org/10.1007/s00330-023-10213-1

[10]LYU Q, TAN J, ZAPADKA M E, et al. Translating radiology reports into plain language using ChatGPT and GPT-4 with prompt learning: results, limitations, and potential. Visual Computing for Industry, Biomedicine, and Art, 2023, 6(1): 9. DOI: https://doi.org/10.1186/s42492-023-00136-5

[11]AYERS J W, POLIAK A, DREDZE M, et al. Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum. JAMA Internal Medicine, 2023, 183(6): 589-596. DOI: https://doi.org/10.1001/jamainternmed.2023.1838

[12]KUCKELMAN I J, WETLEY K, YI P H, et al. Translating musculoskeletal radiology reports into patient-friendly summaries using ChatGPT-4. Skeletal Radiology, 2024. DOI: https://doi.org/10.1007/s00256-024-04599-2

[13]KOPF S, BEAUFILS P, HIRSCHMANN M T, et al. Management of traumatic meniscus tears: the 2019 ESSKA meniscus consensus. Knee Surgery, Sports Traumatology, Arthroscopy, 2020, 28(4): 1177-1194. DOI: https://doi.org/10.1007/s00167-020-05847-3

[14]American Academy of Orthopaedic Surgeons. Management of Anterior Cruciate Ligament Injuries: Evidence-Based Clinical Practice Guideline. Rosemont, IL: American Academy of Orthopaedic Surgeons, 2022.

[15]BEAUFILS P, BECKER R, KOPF S, et al. Surgical management of degenerative meniscus lesions: the 2016 ESSKA meniscus consensus. Knee Surgery, Sports Traumatology, Arthroscopy, 2017, 25(2): 335-346. DOI: https://doi.org/10.1007/s00167-016-4407-4

[16]DeepSeek-AI. DeepSeek-V3 technical report[Z/OL]. arXiv:2412.19437, 2024.

[17]GLM T, ZENG A, XU B, et al. ChatGLM: a family of large language models from GLM-130B to GLM-4 all tools[Z/OL]. arXiv:2406.12793, 2024.

[18]YANG A, YANG B, ZHANG B, et al. Qwen2.5 technical report[Z/OL]. arXiv:2412.15115, 2024.

[19]Team Kimi. Kimi k1.5: scaling reinforcement learning with LLMs[Z/OL]. arXiv:2501.12599, 2025.

[20]MEREDITH S J, RAUER T, CHMIELEWSKI T L, et al. Return to sport after anterior cruciate ligament injury: Panther Symposium ACL Injury Return to Sport Consensus Group. Orthopaedic Journal of Sports Medicine, 2020, 8(6): 2325967120930829. DOI: https://doi.org/10.1177/2325967120930829

[21]LANDIS J R, KOCH G G. The measurement of observer agreement for categorical data. Biometrics, 1977, 33(1): 159-174. DOI: https://doi.org/10.2307/2529310

[22]KOO T K, LI M Y. A guideline of selecting and reporting intraclass correlation coefficients for reliability research. Journal of Chiropractic Medicine, 2016, 15(2): 155-163. DOI: https://doi.org/10.1016/j.jcm.2016.02.012

[23]GUNNING R. The Technique of Clear Writing. New York: McGraw-Hill, 1952.

[24]KINCAID J P, FISHBURNE R P, ROGERS R L, et al. Derivation of New Readability Formulas (Automated Readability Index, Fog Count, and Flesch Reading Ease Formula) for Navy Enlisted Personnel. Millington, TN: Naval Technical Training Command, 1975. DOI: https://doi.org/10.21236/ADA006655

[25]SENTER R J, SMITH E A. Automated readability index. Dayton, OH: Aerospace Medical Research Laboratories, 1967.

[26]COLEMAN M, LIAU T L. A computer readability formula designed for machine scoring. Journal of Applied Psychology, 1975, 60(2): 283-284. DOI: https://doi.org/10.1037/h0076540

[27]DEVLIN J, CHANG M W, LEE K, et al. BERT: pre-training of deep bidirectional transformers for language understanding//Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Stroudsburg, PA: Association for Computational Linguistics, 2019: 4171-4186. DOI: https://doi.org/10.18653/v1/N19-1423

[28]WILCOXON F. Individual comparisons by ranking methods. Biometrics Bulletin, 1945, 1(6): 80-83. DOI: https://doi.org/10.2307/3001968

[29]BENJAMINI Y, HOCHBERG Y. Controlling the false discovery rate: a practical and powerful approach to multiple testing. Journal of the Royal Statistical Society: Series B, 1995, 57(1): 289-300. DOI: https://doi.org/10.1111/j.2517-6161.1995.tb02031.x

[30]KANG H, LI X, FENG M, et al. Applying Radiomics and Deep Learning to Investigations into Seafarers' Health Status. Applied Artificial Intelligence Research, 2025, 1(3). DOI: https://doi.org/10.65455/nqcbkw66

Downloads

Published

2026-09-01

Data Availability Statement

The data that support the findings of this study are available from the corresponding author upon reasonable request.

How to Cite

Chinese Large Language Models for Patient-Friendly Knee MRI Summaries in Sports Injury: A Guideline-Anchored Study of Quality and Robustness. (2026). Applied Artificial Intelligence Research, 2(3), 72-80. https://doi.org/10.65455/2c4jes43