Artifact Evaluation for the Paper: Benchmarking and Understanding Safety Risks in AI Character Platforms
收藏资源简介:
The dataset is the results of the measurement conducted in the paper "Benchmarking and Understanding Safety Risks in AI Character Platforms", accepted to NDSS '26. It contains (i) the metadata of the characters; (ii) the benchmark questions and the corresponding answers obtained from the characters; and (iii) the safety assessment results of the question-answer pairs. We believe that this dataset can serve as a valuable resource for future research, enabling deeper investigations into the safety challenges of AI character platforms and fostering the development of improved safety measures. Columns: - platform: Unique platform identifier - group: Is it from the "Popular Set" or the "Random Set" as described in the paper - bot: Unique character identifier - description: Plain text description of the character, as provided by the platform - tags: Tags associated with the character, as provided by the platform - NSFW: Is the character in NSFW mode (i.e., expected to generate adult content) - question: The benchmark question - answer: The answer to the benchmark question generated by the character - scenario: The opening scenario. At the beginning of the conversation, there will usually be a paragraph that sets the scene and provides background information for the conversation with the character. - question_category: The category of the benchmark question, as provided by the benchmark dataset - judge_category: The safety category in the safety evaluation result, as generated by MD-Judge - judge_score: The safety score in the safety evaluation result, as generated by MD-Judge



