• 1. Department of Thoracic Surgery, West China Hospital, Sichuan University, Chengdu, 610041, P. R. China;
  • 2. West China School of Medicine, Sichuan University, Chengdu, 610041, P. R. China;
  • 3. Institute of Intelligent Computing, University of Electronic Science and Technology of China, Chengdu, 611731, P. R. China;
  • 4. School of Computer Science and Engineering, University of Electronic Science and Technology of China, Chengdu, 611731, P. R. China;
CHEN Nan, Email: dr.chennan@wchscu.cn; LIU Lunxu, Email: lunxu_liu@aliyun.com
Export PDF Favorites Scan Get Citation

Objective  To evaluate the performance of lightweight Chinese large language models (LLMs) in answering specialized lung cancer questions, and to explore the impact of retrieval-augmented generation (RAG) on model performance. Methods  Eleven lightweight Chinese LLMs with parameter sizes ranging from 7B to 32B were included. A lung cancer-specific evaluation dataset consisting of 200 questions [100 A1-type (basic knowledge) and 100 A2-type (clinical case) questions], constructed based on clinical guidelines and thoracic surgery textbooks, was used for assessment. Model performance was evaluated under two conditions (with and without RAG). Accuracy was used to assess model performance, and response latency was recorded to reflect inference efficiency. An accuracy–latency scatter plot was constructed for descriptive analysis of overall model performance. Results  All models successfully completed the evaluation. With the introduction of RAG, the overall average accuracy improved from 61.68% to 76.36%. Smaller models demonstrated the most significant improvement (e.g., the accuracy of DeepSeek-7B increased from 32.50% to 60.00%, P<0.001). The average response latency increased from 12.58 s to 13.80 s. The Qwen3 series showed the best overall performance, and Qwen3-32B achieved the highest accuracy under both conditions (76.50% and 84.00%, respectively). After RAG integration, performance differences among model families were markedly reduced. Based on the accuracy-latency trade-off, Qwen3-32B achieved the best balance between accuracy and response latency under the baseline condition, whereas Qwen3-14B demonstrated superior overall performance in terms of accuracy, latency, and computational cost after RAG integration. Conclusion  The integration of RAG technology improves the ability of lightweight Chinese LLMs to answer specialized lung cancer questions. Under the dual practical constraints of limited computational resources and medical data security requirements, the "lightweight model+RAG" technical framework may represent a promising deployment solution.

Copyright ? the editorial department of Chinese Journal of Clinical Thoracic and Cardiovascular Surgery of West China Medical Publisher. All rights reserved