Chinese-SimpleQA
收藏资源简介:
Chinese SimpleQA是一个全面的中文基准,用于评估语言模型回答简短问题的事实性能力。它主要具有五个特性:中文、多样性、高质量、静态和易于评估。数据集涵盖了6个主要主题,包括99个细分的子主题,从人文到科学和工程。数据集包含3000个高质量问题,旨在帮助开发者更深入地理解其模型在中文领域的事实正确性,并为算法研究提供重要基础。
Chinese SimpleQA is a comprehensive Chinese benchmark for evaluating the factual capability of language models when answering short questions. It has five core characteristics: Chinese-language support, diversity, high quality, static nature, and ease of evaluation. The dataset covers 6 major themes, including 99 subdivided subtopics ranging from humanities to science and engineering. It contains 3,000 high-quality questions, designed to help developers gain a deeper understanding of their models' factual correctness in the Chinese language domain, and provide a crucial foundation for algorithmic research.
Chinese SimpleQA 数据集概述
基本信息
- 许可证: cc-by-nc-sa-4.0
- 任务类别: 问答
- 语言: 中文
- 数据集名称: Chinese SimpleQA
- 数据规模: 10K<n<100K
数据集简介
- 目标: 评估语言模型回答简短问题的真实性能力。
- 主要特点:
- 中文: 专注于中文语言,全面评估现有大型语言模型(LLMs)在中文方面的真实性能力。
- 多样性: 涵盖6个主要主题,包括“中国文化”、“人文”、“工程、技术与应用科学”、“生活、艺术与文化”、“社会”和“自然科学”,共计99个细粒度子主题。
- 高质量: 通过全面严格的质量控制流程,确保数据集的质量和准确性。
- 静态: 所有参考答案不会随时间变化,保持数据集的常青特性。
- 易于评估: 问题和答案非常简短,可以通过现有的LLMs(如OpenAI API)快速运行评分程序。
数据集内容
- 主题覆盖: 6个主要主题,99个细粒度子主题。
- 问题数量: 3000个高质量问题。
引用
@misc{he2024chinesesimpleqachinesefactuality, title={Chinese SimpleQA: A Chinese Factuality Evaluation for Large Language Models}, author={Yancheng He and Shilong Li and Jiaheng Liu and Yingshui Tan and Weixun Wang and Hui Huang and Xingyuan Bu and Hangyu Guo and Chengwei Hu and Boren Zheng and Zhuoran Lin and Xuepeng Liu and Dekai Sun and Shirong Lin and Zhicheng Zheng and Xiaoyong Zhu and Wenbo Su and Bo Zheng}, year={2024}, eprint={2411.07140}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2411.07140}, }




