Chinese SafetyQA
收藏资源简介:
Chinese SafetyQA是一个专门为评估大型语言模型(LLMs)在安全知识方面的短形式事实性基准数据集。该数据集由中国阿里巴巴集团未来实验室开发,涵盖了七个主要的安全领域,包括法律框架、伦理标准等。数据集包含2000个高质量的安全示例,分为问答(QA)和多选题(MCQ)两种格式,旨在全面评估LLMs在处理安全相关问题时的准确性和稳健性。数据集的创建过程包括从搜索引擎和官方网站收集种子示例,通过GPT4o进行数据增强和QA对生成,并经过多轮验证和人工标注确保数据质量。该数据集主要用于评估和提升LLMs在法律、政策和伦理等领域的安全知识能力,旨在解决模型在处理安全问题时可能产生的幻觉和错误。
Chinese SafetyQA is a short-form factual benchmark dataset dedicated to evaluating large language models (LLMs) on their safety knowledge. Developed by the Future Lab of Alibaba Group in China, this dataset covers seven major safety domains including legal frameworks and ethical standards. It contains 2,000 high-quality safety instances available in two formats: question-answering (QA) and multiple-choice question (MCQ), aiming to comprehensively assess the accuracy and robustness of LLMs when addressing safety-related issues. The dataset’s development pipeline includes collecting seed samples from search engines and official websites, performing data augmentation and QA pair generation via GPT-4o, and undergoing multi-round validation and manual annotation to ensure data quality. This dataset is primarily used to evaluate and enhance the safety knowledge capabilities of LLMs across fields such as law, policy and ethics, with the goal of mitigating hallucinations and erroneous outputs that models may produce when handling safety-related problems.

- 1Chinese SafetyQA: A Safety Short-form Factuality Benchmark for Large Language Models阿里巴巴集团未来实验室 · 2024年



