CIF-Bench
收藏资源简介:
CIF-Bench是由曼彻斯特大学等机构创建的中文指令遵循基准,用于评估大型语言模型对中文语言的零样本泛化能力。该数据集包含150个任务和45,000个输入输出对,由母语为中文的标注者开发,测试复杂推理和中文文化细微差别。为减少评估偏差,数据集的一半公开发布,另一半保持私有,并引入多样化指令以最小化分数差异。CIF-Bench旨在揭示当前LLMs处理中文任务的局限性,推动开发更具文化敏感性和语言多样性的模型。
CIF-Bench is a Chinese instruction-following benchmark developed by the University of Manchester and other institutions, designed to evaluate the zero-shot generalization capability of large language models (LLMs) on Chinese language tasks. This dataset includes 150 tasks and 45,000 input-output pairs, developed by native Chinese annotators, and is intended to assess complex reasoning skills and subtle Chinese cultural nuances. To reduce evaluation bias, half of the dataset is publicly released while the other half remains private, and diverse instructions are introduced to minimize score discrepancies. CIF-Bench aims to reveal the limitations of current LLMs in handling Chinese-related tasks, and promote the development of models with greater cultural sensitivity and linguistic diversity.




