• 1. Department of Thoracic Surgery, West China Hospital, Sichuan University, Chengdu, 610041, P. R. China;
  • 2. West China School of Medicine, Sichuan University, Chengdu, 610041, P. R. China;
  • 3. Institute of Intelligent Computing, University of Electronic Science and Technology of China, Chengdu, 611731, P. R. China;
  • 4. School of Computer Science and Engineering, University of Electronic Science and Technology of China, Chengdu, 611731, P. R. China;
CHEN Nan, Email: dr.chennan@wchscu.cn; LIU Lunxu, Email: lunxu_liu@aliyun.com
Export PDF Favorites Scan Get Citation

Objective  To evaluate the performance of lightweight Chinese large language models (LLMs) in answering specialized lung cancer questions, and to explore the impact of retrieval-augmented generation (RAG) on model performance. Methods  Eleven lightweight Chinese LLMs with parameter sizes ranging from 7B to 32B were included. A lung cancer-specific evaluation dataset consisting of 200 questions [100 A1-type (basic knowledge) and 100 A2-type (clinical case) questions], constructed based on clinical guidelines and thoracic surgery textbooks, was used for assessment. Model performance was evaluated under two conditions (with and without RAG). Accuracy was used to assess model performance, and response latency was recorded to reflect inference efficiency. An accuracy–latency scatter plot was constructed for descriptive analysis of overall model performance. Results  All models successfully completed the evaluation. With the introduction of RAG, the overall average accuracy improved from 61.68% to 76.36%. Smaller models demonstrated the most significant improvement (e.g., the accuracy of DeepSeek-7B increased from 32.50% to 60.00%, P<0.001). The average response latency increased from 12.58 s to 13.80 s. The Qwen3 series showed the best overall performance, and Qwen3-32B achieved the highest accuracy under both conditions (76.50% and 84.00%, respectively). After RAG integration, performance differences among model families were markedly reduced. Based on the accuracy-latency trade-off, Qwen3-32B achieved the best balance between accuracy and response latency under the baseline condition, whereas Qwen3-14B demonstrated superior overall performance in terms of accuracy, latency, and computational cost after RAG integration. Conclusion  The integration of RAG technology improves the ability of lightweight Chinese LLMs to answer specialized lung cancer questions. Under the dual practical constraints of limited computational resources and medical data security requirements, the "lightweight model+RAG" technical framework may represent a promising deployment solution.

Citation: REN Zhizhen, LI Yile, MA Yizhuo, CHEN Qizhi, LI Jili, YANG Siyi, ZHANG Jianhao, QIN Ke, PU Qiang, CHEN Nan, LIU Lunxu. Performance evaluation of lightweight Chinese large language models integrated with retrieval-augmented generation technology in answering specialized lung cancer questions. Chinese Journal of Clinical Thoracic and Cardiovascular Surgery, 2026, 33(9): 1403-1411. doi: 10.7507/1007-4848.202604095 Copy

Copyright ? the editorial department of Chinese Journal of Clinical Thoracic and Cardiovascular Surgery of West China Medical Publisher. All rights reserved

  • Previous Article

    Advanced research on respiratory rehabilitation for patients with postoperative pulmonarycomplications after aortic dissection surgery
  • Next Article

    Preoperative nebulized indocyanine green-assisted thoracoscopic anatomical lesion resection for congenital pulmonary airway malformations in children: A retrospective cohort study