CentificAIResearch/trust-and-safety-multiturn-evaluation-dataset
收藏资源简介:
该数据集旨在评估基于大型语言模型(LLM)的评分模型在安全相关对话中的性能。数据集包含在多轮安全评估期间生成的模型响应,以及指示这些响应是否符合安全策略的参考标签。该基准测试评估评分模型在不同安全类别中准确识别安全和不安全模型行为的能力。数据集创建基于从多轮安全对话中收集的模型响应,目标模型的响应根据预定义的安全策略进行评估并分配参考标签。评分模型的任务是评估相同的响应,从而比较评分输出与参考标签。数据集涵盖多个安全领域,包括但不限于:自我伤害、暴力、化学与生物风险、仇恨与骚扰、性内容、儿童安全和错误信息。数据结构示例包含对话ID、轮次ID、策略名称、模型响应、参考标签、评分标签和评分理由。
This dataset is designed to evaluate the performance of LLM-based graders on safety-related conversations. The dataset consists of model responses generated during multi-turn safety evaluations along with reference labels indicating whether the responses comply with safety policies. This benchmark evaluates a graders ability to accurately identify safe and unsafe model behavior across different safety categories. The benchmark was created using model responses collected from multi-turn safety conversations. Responses from target models were evaluated against predefined safety policies and assigned reference labels. Grader models were then tasked with assessing the same responses, enabling comparison between grader outputs and reference labels. The benchmark includes conversations across multiple safety domains, including but not limited to: Self-Harm, Violence, Chemical and Biological Risks, Hate and Harassment, Sexual Content, Child Safety, and Misinformation. The dataset structure includes fields such as conversation_id, turn_id, policy_name, model_response, reference_label, grader_label, and grader_reasoning.




