• 1. West China Hospital, Sichuan University, Chengdu, Sichuan 610041, P. R. China;
  • 2. Department of Postgraduate Students, West China Hospital, Sichuan University, Chengdu, Sichuan 610041, P. R. China;
  • 3. West China Medical Publishers, West China Hospital, Sichuan University, Chengdu, Sichuan 610041, P. R. China;
  • 4. Colorectal Cancer Center, West China Hospital, Sichuan University, Chengdu, Sichuan 610041, P. R. China;
ZHOU Yuwen, Email: drzhouyuwen@163.com
Export PDF Favorites Scan Get Citation

Objective  To systematically evaluate the consistency between large language models (LLMs) and human raters in the assessment of medical students’ knowledge examinations, clinical documentation, and behavioral performance. Methods  PubMed, Web of Science, Embase, ERIC, China National Knowledge Infrastructure, and WanFang Data were searched for original research articles published between January 2020 and March 2026 regarding the agreement between LLMs and human scoring in medical student assessments. Meta-analysis was performed using R 5.3.0 software. Results  A total of 16 studies were included, yielding 39 independent effect sizes, comprising 17 intraclass correlation coefficients (ICC) and 22 Cohen’s Kappa data points. The assessment tasks spanned three categories: knowledge examinations, clinical documentation, and behavioral assessment, involving various mainstream LLMs such as GPT-4. Meta-analysis revealed a pooled ICC of 0.74 [95% confidence interval (CI) (0.51, 0.87), P<0.001] and a pooled Cohen’s Kappa of 0.52 [95%CI (0.38, 0.64), P<0.001], indicating overall moderate-to-high consistency. Significant heterogeneity was observed among the included studies, primarily driven by task types and task difficulty. Conclusions  Constrained by factors such as task difficulty, LLMs currently serve only as adjunctive tools to human evaluation, despite demonstrating a potential for consistency comparable to human scoring in medical education assessments. The future intelligent transformation of medical education evaluation must operate within a clinician-in-the-loop framework to achieve a balance between assessment efficiency and the rigor of medical logic.

Citation: LI Huayu, ZHANG Qin, XU Qian, QIU Meng, ZHOU Yuwen. Agreement between large language models and human scoring in medical student assessments: a meta-analysis. West China Medical Journal, 2026, 41(6): 983-989. doi: 10.7507/1002-0179.202603259 Copy

Copyright ? the editorial department of West China Medical Journal of West China Medical Publisher. All rights reserved

  • Previous Article

    Exploration of the early training system for medical-engineering interdisciplinary postgraduates in the context of “intelligent digital rehabilitation”
  • Next Article

    Application effect of DeepSeek-empowered virtual simulation teaching in neonatal resuscitation clinical training