Objective To systematically evaluate the consistency between large language models (LLMs) and human raters in the assessment of medical students’ knowledge examinations, clinical documentation, and behavioral performance. Methods PubMed, Web of Science, Embase, ERIC, China National Knowledge Infrastructure, and WanFang Data were searched for original research articles published between January 2020 and March 2026 regarding the agreement between LLMs and human scoring in medical student assessments. Meta-analysis was performed using R 5.3.0 software. Results A total of 16 studies were included, yielding 39 independent effect sizes, comprising 17 intraclass correlation coefficients (ICC) and 22 Cohen’s Kappa data points. The assessment tasks spanned three categories: knowledge examinations, clinical documentation, and behavioral assessment, involving various mainstream LLMs such as GPT-4. Meta-analysis revealed a pooled ICC of 0.74 [95% confidence interval (CI) (0.51, 0.87), P<0.001] and a pooled Cohen’s Kappa of 0.52 [95%CI (0.38, 0.64), P<0.001], indicating overall moderate-to-high consistency. Significant heterogeneity was observed among the included studies, primarily driven by task types and task difficulty. Conclusions Constrained by factors such as task difficulty, LLMs currently serve only as adjunctive tools to human evaluation, despite demonstrating a potential for consistency comparable to human scoring in medical education assessments. The future intelligent transformation of medical education evaluation must operate within a clinician-in-the-loop framework to achieve a balance between assessment efficiency and the rigor of medical logic.