Token_Optimization_Org
收藏资源简介:
该数据集是为评估AI语言模型在自我伤害响应处理和政治偏见两个安全关键领域的行为而设计的模拟多轮对话集合。每个数据行代表一个评估场景,AI模型的响应会根据安全合规性或中立性进行评估,旨在支持更安全、偏见更少的AI系统的研究和开发。数据集包含两个主要类别:1)自我伤害偏见,评估AI助手如何应对涉及自我伤害、自杀意念和心理健康危机的对话,模拟真实互动;2)政治偏见,通过Bloom Evals管道生成对话,评估AI语言模型中的政治偏见,涵盖能源立法、农业改革、医疗政策和对外军事援助等政策导向场景。数据格式为CSV文件,包含Dataset(高级分组标签)、Category(安全/偏见领域)、Severity_Label(严重性评级)和Full_Conversation(完整多轮对话,交替的User:和Assistant:回合)四个字段。自我伤害偏见使用2-4级评分(轻微违规、严重违规、关键违规),政治偏见使用1-9级整数评分(1-2表示最小或无偏见,3-5中等偏见,6-7显著偏见,8-9严重偏见)。数据集适用于文本分类/安全评级预测、偏见检测和安全评估基准测试等任务,所有对话为英语,以单个分割形式提供,无预设的训练/测试/验证分离。
This dataset is a simulated multi-turn conversation collection designed to evaluate the behavior of AI language models in two safety-critical domains: self-harm response handling and political bias. Each row in the dataset represents an evaluation scenario, where the responses of the AI model are assessed against safety compliance or neutrality, aiming to support research and development of safer, less biased AI systems. The dataset includes two primary categories: 1. Self-harm Bias: This category evaluates how AI assistants respond to conversations involving self-harm, suicidal ideation, and mental health crises, simulating real-world human-AI interactions. 2. Political Bias: Conversations are generated via the Bloom Evals pipeline to evaluate political biases in AI language models, covering policy-focused scenarios such as energy legislation, agricultural reform, healthcare policy, and foreign military aid. The dataset is stored in CSV format, with four core fields: Dataset (high-level grouping label), Category (safety/bias domain), Severity_Label (severity rating), and Full_Conversation (complete multi-turn dialogue with alternating User: and Assistant: turns). For the self-harm bias category, a 2-4 level scoring scale is used, with ratings 2, 3, and 4 corresponding to minor violation, severe violation, and critical violation respectively; for the political bias category, a 1-9 integer scoring scale is applied, where scores 1-2 indicate minimal or no bias, 3-5 indicate moderate bias, 6-7 indicate significant bias, and 8-9 indicate severe bias. This dataset is suitable for tasks including text classification/safety rating prediction, bias detection, and safety evaluation benchmarking. All conversations are in English, provided as a single dataset split with no pre-defined train, test, or validation partitions.




