| 1. |
Holmboe ES, Sherbino J, Long DM, et al. The role of assessment in competency-based medical education. Med Teach, 2010, 32(8): 676-682.
|
| 2. |
Yeates P, O’Neill P, Mann K, et al. Seeing the same thing differently. Adv Health Sci Educ, 2013, 18(3): 325-341.
|
| 3. |
Thirunavukarasu AJ, Ting DSJ, Elangovan K, et al. Large language models in medicine. Nat Med, 2023, 29(8): 1930-1940.
|
| 4. |
Bolgova O, Ganguly P, Ikram MF, et al. Evaluating large language models as graders of medical short answer questions: a comparative analysis with expert human graders. Med Educ Online, 2025, 30(1): 2550751.
|
| 5. |
Singhal K, Azizi S, Tu T, et al. Large language models encode clinical knowledge. Nature, 2023, 620(7972): 172-180.
|
| 6. |
Sim SZY, Chen T. Critique of impure reason: unveiling the reasoning behaviour of medical large language models. Elife, 2025, 14: e106187.
|
| 7. |
Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ, 2021, 372: n71.
|
| 8. |
Lucas N, Macaskill P, Irwig L, et al. The reliability of a quality appraisal tool for studies of diagnostic reliability (QAREL). BMC Med Res Methodol, 2013, 13: 111.
|
| 9. |
Grévisse C. LLM-based automatic short answer grading in undergraduate medical education. BMC Med Educ, 2024, 24: 1060.
|
| 10. |
Quah B, Zheng L, Sng TJH, et al. Reliability of ChatGPT in automated essay scoring for dental undergraduate examinations. BMC Med Educ, 2024, 24(1): 962.
|
| 11. |
Bond WF, Zhou J, Bhat S, et al. Automated patient note grading: examining scoring reliability and feasibility. Acad Med, 2023, 98(Suppl 3): S90-S97.
|
| 12. |
Holcomb MJ, Kang S, Shakur A, et al. Zero-shot multimodal question answering for assessment of medical student OSCE physical exam videos. medRxiv, 2024: 2024.06.05.24308467.
|
| 13. |
Seneviratne HMTW, Manathunga SS. Artificial intelligence assisted automated short answer question scoring tool shows high correlation with human examiner markings. BMC Med Educ, 2025, 25(1): 1146.
|
| 14. |
Rajan A, Alexander SM, Shenvi CL. Can AI grade like a professor? Comparing artificial intelligence and faculty scoring of medical student short-answer clinical reasoning exams. Adv Health Sci Educ Theory Pract, 2026, 31(2): 607-617.
|
| 15. |
Wang L, Mao Y, Wang L, et al. Suitability of GPT-4o as an evaluator of cardiopulmonary resuscitation skills examinations. Resuscitation, 2024, 204: 110404.
|
| 16. |
Jade T, Yartsev A. ChatGPT for automated grading of short-answer questions in mechanical ventilation examinations. Focus Health Prof Educ, 2026, 27(1): 90-104.
|
| 17. |
Yokose M, Hirosawa T, Sakamoto T, et al. The validity of generative artificial intelligence in evaluating medical students in objective structured clinical examination: experimental study. JMIR Form Res, 2025, 9: e79465.
|
| 18. |
Bentegeac R, Florens N, Maanaoui M, et al. ECOSBot: a multicenter validation pilot study of a generative AI tool for OSCE-based nephrology training. Clin Kidney J, 2025, 18(10): sfaf308.
|
| 19. |
Niu S, Liu X, Huang L, et al. DeepSeek-R1 for automated scoring in radiology residency examinations: an agreement and test-retest reliability study. BMC Med Educ, 2025, 25(1): 1581.
|
| 20. |
Sreedhar R, Chang L, Gangopadhyaya A, et al. Comparing scoring consistency of large language models with faculty for formative assessments in medical education. J Gen Intern Med, 2025, 40(1): 127-134.
|
| 21. |
Luordo D, Arrese MT, Calvo CT, et al. Application of artificial intelligence as an aid for the correction of the objective structured clinical examination (OSCE). Appl Sci, 2025, 15(3): 1153.
|
| 22. |
Atsukawa N, Tatekawa H, Oura T, et al. Evaluation of radiology residents’ reporting skills using large language models: an observational study. Jpn J Radiol, 2025, 43(7): 1204-1212.
|
| 23. |
Shakur AH, Holcomb MJ, Hein D, et al. Large language models for medical OSCE assessment: a novel approach to transcript analysis. arXiv, 2024: 2410.12858.
|
| 24. |
Hassanein FEA, Hussein RR, Ahmed Y, et al. Calibration of AI large language models with human subject matter experts for grading of clinical short-answer responses in dental education. BMC Oral Health, 2026, 26: 286.
|
| 25. |
Takahashi H, Shikino K, Kondo T, et al. AI- vs human-based assessment of medical interview transcripts in a generative AI-simulated patient system: cross-sectional validation study. JMIR Med Educ, 2026, 12(1): e81673.
|
| 26. |
Berger A, Khanna S, Berghaus D, et al. Reasoning LLMs in the medical domain: a literature survey. arXiv, 2025: 2508.19097v1.
|
| 27. |
Prenosil GA, Weitzel TK, Bello SC, et al. Neuro-symbolic AI for auditable cognitive information extraction from medical reports. Commun Med (Lond), 2025, 5(1): 491.
|
| 28. |
Mishra V, Lurie Y, Mark S. Accuracy of LLMs in medical education: Evidence from a concordance test with medical teacher. BMC Med Educ, 2025, 25: 443.
|
| 29. |
Gong EJ, Bang CS, Lee JJ, et al. Knowledge-practice performance gap in clinical large language models: systematic review of 39 benchmarks. J Med Internet Res, 2025, 27: e84120.
|
| 30. |
Mo?ll B, Sand Aronsson F, Akbar S. Medical reasoning in LLMs: an in-depth analysis of DeepSeek R1. Front Artif Intell, 2025, 8: 1616145.
|
| 31. |
Kang S, Holcomb MJ, Shakur AH, et al. Automated assessment of OSCE physical exams using multimodal AI. medRxiv, 2026: 2026.01.09.26343786.
|
| 32. |
Boscardin CK, Gin B, Golde PB, et al. ChatGPT and generative artificial intelligence for medical education: potential impact and opportunity. Acad Med, 2024, 99(1): 22-27.
|
| 33. |
Zack T, Lehman E, Suzgun M, et al. Assessing the potential of GPT-4 to perpetuate racial and gender biases in health care: a model evaluation study. Lancet Digit Health, 2024, 6(1): e12-e22.
|
| 34. |
Gu Y, Chen X, Shan G, et al. Development and validation of a deep learning-based emergency triage model: a feasibility and effectiveness study. BMC Emerg Med, 2026, 26: 59.
|
| 35. |
Yap WS, Saw PS, Yeap LL, et al. Comparison between GPT-4 and human raters in grading pharmacy students’ exam responses in Malaysia: a cross-sectional study. J Educ Eval Health Prof, 2025, 22: 20.
|
| 36. |
Roveta A, Castello LM, Massarino C, et al. Artificial intelligence in medical education: a narrative review on implementation, evaluation, and methodological challenges. AI, 2025, 6(9): 227.
|
| 37. |
Hoebers FJP, Wee L, Likitlersuang J, et al. Artificial intelligence research in radiation oncology: a practical guide for the clinician on concepts and methods. BJR Open, 2024, 6(1): tzae039.
|
| 38. |
Selected abstracts from the SIIM 2025 annual meeting of the society for imaging informatics in medicine (SIIM). J Imaging Inform Med, 2025, 38(Suppl 1): 1-49.
|
| 39. |
Hasan M, Hossain A, Sayem FH, et al. CLIN-LLM: a safety-constrained hybrid framework for clinical diagnosis and treatment generation. arXiv, 2025: 2510.22609v1.
|
| 40. |
Kung TH, Cheatham M, Medenilla A, et al. Performance of ChatGPT on USMLE: potential for AI-assisted medical education using large language models. PLOS Digit Health, 2023, 2(2): e0000198.
|
| 41. |
Thompson RAM, Shah YB, Aguirre F, et al. Artificial intelligence use in medical education: best practices and future directions. Curr Urol Rep, 2025, 26(1): 45.
|
| 42. |
Masters K. Ethical use of artificial intelligence in health professions education: AMEE guide no.158. Med Teach, 2023, 45(6): 574-584.
|