Safe-Child-LLM
收藏资源简介:
Safe-Child-LLM是一个专门为评估大型语言模型(LLMs)在儿童和青少年交互中的安全性而设计的基准和数据集。该数据集包含200个对抗性提示,这些提示是从公开的红色团队语料库(如SG-Bench、HarmBench)中精心挑选和修改的,并带有人工标注的标签,用于评估越狱的成功率和一个标准化的0-5道德拒绝等级。该数据集旨在通过包含真实场景的提示,反映儿童(7-12岁)和青少年(13-17岁)在不同发展阶段的需求。为了促进透明度和协作推进伦理AI开发,Safe-Child-LLM的基准数据集和评估代码库已公开发布,供AI安全社区使用和改进。
Safe-Child-LLM is a benchmark and dataset specifically designed to evaluate the safety of large language models (LLMs) during interactions with children and adolescents. This dataset contains 200 adversarial prompts, which are carefully selected and modified from public red team corpora such as SG-Bench and HarmBench, with manually annotated labels for evaluating jailbreak success rates and a standardized 0-5 moral refusal rating. This dataset aims to reflect the needs of children (aged 7-12) and adolescents (aged 13-17) at different developmental stages by including prompts based on real-world scenarios. To promote transparency and facilitate collaborative progress in ethical AI development, the benchmark dataset and evaluation codebase of Safe-Child-LLM have been publicly released for the AI safety community to utilize and improve upon.




