Dataset for Reliability of Large Language Models as Evaluators of Traditional Chinese Medicine Terminology
收藏资源简介:
This dataset supports the study “Reliability of Large Language Models as Evaluators of Traditional Chinese Medicine Terminology”. It contains two files. The first comprises the complete eligible corpus of 376 terminology contrasts derived from the WHO International Standard Terminologies on Traditional Chinese Medicine published in 2022. Each record includes the Chinese term, WHO-listed English term, listed synonym, and WHO English definition used in the evaluation task. The second contains the response-level data generated through the evaluation of these 376 terminology contrasts using four large language model configurations under two candidate-order conditions and three independent calls per order. It includes model and request metadata, raw responses, parsing information, criterion-level and overall decisions, mapped decision variables, justifications, completion records, and format-compliance information. The complete evaluation prompt and analytical procedures are reported in the associated manuscript.



