遇见数据集

giskardai/StereoTales

收藏
Hugging Face2026-05-04 更新2026-06-14 收录
官方服务:

资源简介:

这是一个多语言评估数据集,用于探测大型语言模型(LLM)在故事生成中的人口统计偏见。每个样本指导模型生成一个约200词的故事,故事中的角色具有给定的人口统计属性值(如年龄、性别、民族、宗教、残疾状况、移民状况等),并置于特定的生活场景中,旨在揭示生成叙事中的社会经济和人口统计偏见。数据集支持多种语言(包括英语、阿拉伯语、荷兰语、西班牙语、法语、印地语、意大利语、葡萄牙语、乌克兰语和中文),按语言组织为独立的配置,便于用户加载特定语言子集。所有语言配置共享相同的结构(相同的属性、属性值、场景和提示模板),仅自然语言内容经过翻译和人工审核,从而支持跨语言模型行为的直接比较。此外,数据集还包括额外的英语配置(不含场景的提示)、生成的故事样本(包含模型输出和属性提取结果)、以及分析和研究工件(如人类评估、模型自评估和关联列表的Parquet表)。数据集遵循Flare的Sample模式,每个样本包含ID、模块、任务、语言、生成内容、元数据和评估字段。

A multilingual evaluation dataset for probing demographic biases in LLM story generation. Each sample instructs a model to write a ~200-word story about a character carrying a given demographic attribute value (age, gender, ethnicity, religion, disability status, immigration status, ...) placed into a specific life scenario, with the goal of surfacing socio-economic and demographic biases in the generated narratives. The dataset is organized as one config per language, allowing users to load specific language subsets independently. All language configs share the same structure: identical attributes, attribute values, scenarios, and prompt templates — only the natural-language content is translated and human-reviewed, enabling direct cross-lingual comparison of model behavior. It includes additional English-only configs with scenario-free prompts, generated story samples with model outputs and attribute extraction results, and analysis artifacts such as human evaluation, model self-evaluation, and associations listings in Parquet tables. The dataset follows the Flare Sample schema, with each sample containing fields like ID, module, task, language, generations, metadata, and evaluation.

提供机构:
giskardai
二维码
社区交流群
二维码
科研交流群
商业服务