In recent years, the rapid advancement of artificial intelligence has profoundly affected societal structures and a broad range of sectors, including healthcare, and is reshaping modes of production and everyday life. Across the evolution of AI technologies, high-quality datasets provide the essential foundation for model training and iterative improvement. Nevertheless, the construction of datasets for AI training remains challenged by ambiguous objective definition and insufficient core technical support, which together hinder the development of high-quality AI datasets. To address these challenges, the China Academy of Information and Communications Technology and Tsinghua University, together with other institutions, jointly released the Guidelines for the Construction of High-Quality Datasets for Artificial Intelligence (hereinafter, the Guidelines), proposing an engineering-oriented, full-lifecycle framework that encompasses data collection, governance, annotation, quality inspection, and operational management. Centered on the Guidelines, this article reviews and interprets their release context, key definitions, and recommended implementation pathways, with the aim of providing a structured reference for researchers seeking to systematically advance high-quality AI dataset development.
Citation: WU Mengyao, LIAO Mingyu, GUO Jin, ZHENG Mengyuan, ZHAO Peng, XIONG Yiquan, SUN Xin, TAN Jing. Guidelines for the construction of high-quality datasets for artificial intelligence: an interpretation. Chinese Journal of Evidence-Based Medicine, 2026, 26(6): 723-728. doi: 10.7507/1672-2531.202601213 Copy
Copyright ? the editorial department of Chinese Journal of Evidence-Based Medicine of West China Medical Publisher. All rights reserved

