| 1. |
|
| 2. |
Zhao WX, Zhou K, Li J, et al. A survey of large language models. arXiv preprint arXiv, 2023, 230318223.
|
| 3. |
Bai Y, Geng X, Mangalam K, et al. Sequential modeling enables scalable learning for large vision models. IEEE/CVF Conf Com Vis, 2023: 22861-22872.
|
| 4. |
Wang Z, Xia X, Chen R, et al. LaVin-DiT: large vision diffusion transformer. IEEE/CVF Conf Com Vis Pat Rec, 2024: 20060-20070.
|
| 5. |
|
| 6. |
|
| 7. |
Wang L, Lyu C, Ji T, et al. Document-level machine translation with large language models. arXiv preprint arXiv, 2023: 230402210.
|
| 8. |
|
| 9. |
|
| 10. |
|
| 11. |
|
| 12. |
|
| 13. |
|
| 14. |
|
| 15. |
|
| 16. |
|
| 17. |
|
| 18. |
|
| 19. |
|
| 20. |
|
| 21. |
周青青. 基于檢索增強生成的胃病智能問答系統設計與實現. 合肥: 安徽中醫藥大學, 2025.
|
| 22. |
|
| 23. |
|
| 24. |
|
| 25. |
|
| 26. |
|
| 27. |
|
| 28. |
Yamada H, Matsumoto Y. Statistical dependency analysis with support vector machines. Proceed Eighth Intern Conferen Pars Tech, 2003: 195-206.
|
| 29. |
Eisner J. An empirical comparison of probability models for dependency grammar. arXiv preprint arXiv, 1997: 9706004.
|
| 30. |
Nivre J, Scholz M. Deterministic dependency parsing of English text. COLING 2004: Proceedings of the 20th International Conference on Computational Linguistics, 2004.
|
| 31. |
|
| 32. |
|
| 33. |
|
| 34. |
Och FJ. Minimum error rate training in statistical machine translation. Proceedings of the 41st Annual Meeting of the Association for Computational Linguistics, 2003: 160-167.
|
| 35. |
Papineni K, Roukos S, Ward T, et al. BLEU: a method for automatic evaluation of machine translation. Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, 2002: 311-318.
|
| 36. |
Doddington G. Automatic evaluation of machine translation quality using n-gram co-occurrence statistics. Proceedings of the second international conference on Human Language Technology Research, 2002: 138-145.
|
| 37. |
Vedantam R, Zitnick CL, Parikh D. CIDEr: consensus-based image description evaluation. 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015: 4566-4575.
|
| 38. |
Lin CY. ROUGE: a package for automatic evaluation of summaries. Text Summarization Branches Out, 2004: 74-81.
|
| 39. |
Satanjeev B. METEOR: an automatic metric for MT evaluation with improved correlation with human judgments. Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, 2005: 65-72.
|
| 40. |
Sellam T, Das D, Parikh AP. BLEURT: learning robust metrics for text generation. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020: 7881-7892.
|
| 41. |
Zhang T, Kishore V, Wu F, et al. BERTScore: evaluating text generation with BERT. arXiv preprint arXiv, 2019: 190409675.
|
| 42. |
Zhao W, Peyrard M, Liu F, et al. MoverScore: text generation evaluating with contextualized embeddings and earth mover distance. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), 2019: 563-578.
|
| 43. |
Turian JP, Shea L, Melamed ID. Evaluation of machine translation and its evaluation. Proceedings of Machine Translation Summit IX: Papers, 2003.
|
| 44. |
|
| 45. |
Buckley C, Voorhees EM. Evaluating evaluation measure stability. Proceedings of the 23rd Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, 2000: 33-40.
|
| 46. |
Voorhees E. The TREC-8 question answering track report. Text Retrieval Conference, 1999.
|
| 47. |
|
| 48. |
|
| 49. |
|
| 50. |
|
| 51. |
|
| 52. |
|
| 53. |
|
| 54. |
|
| 55. |
|
| 56. |
|
| 57. |
|
| 58. |
|
| 59. |
|
| 60. |
|
| 61. |
|
| 62. |
王興蒙, 戴國華, 高武霖, 等. PROBAST+AI: 更新基于回歸或人工智能方法的預測模型的質量、偏倚風險及適用性系統評價工具—中文解讀. 中國胸心血管外科臨床雜志, 2025: 1-10.
|
| 63. |
|
| 64. |
|
| 65. |
|
| 66. |
|
| 67. |
|
| 68. |
|
| 69. |
|
| 70. |
|
| 71. |
|
| 72. |
|
| 73. |
|
| 74. |
|
| 75. |
|
| 76. |
|
| 77. |
|
| 78. |
|
| 79. |
Heus P, Reitsma JB, Collins GS, et al. Transparent reporting of multivariable prediction models in journal and conference abstracts: TRIPOD for abstracts. Ann Intern Med, 2020: M20-0193.
|
| 80. |
|
| 81. |
|
| 82. |
|
| 83. |
|
| 84. |
李澤宇, 詹正哲, 程嘉儀, 等. 基于人工智能的臨床預測模型研究報告規范(TRIPOD+AI)中文解讀. 中國循證醫學雜志, 2025, 25(3): 339-343.
|
| 85. |
|
| 86. |
|
| 87. |
|
| 88. |
|
| 89. |
|
| 90. |
|
| 91. |
Arias-Duart A, Martin-Torres PA, Hinjos D, et al. Automatic evaluation of healthcare LLMs beyond question-answering. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies, 2025: 108-130.
|
| 92. |
|
| 93. |
|
| 94. |
Kulkarni A, Zhang Y, Moniz JRA, et al. Evaluating evaluation metrics - the mirage of hallucination detection. arXiv preprint arXiv, 2025: 250418114.
|
| 95. |
|