CohereLabs/CultureMarkers
收藏资源简介:
该数据集名为The Culture Funnel: You Cant Align What isnt in the Data,包含约560万个经过文化标签标注的样本,旨在帮助研究人员研究并缓解大型语言模型(LLM)管道中的文化数据漏斗问题。数据集采用多维标签框架,识别预训练、微调、对齐和推理数据集中的文化信号、领域、地理位置和任务专业化。标签包括文化维度(如知识、偏好、动态、偏见、一般文化和无文化)、领域(如数学、人文、对话、社会科学)、任务意图、地理位置以及语言识别。这些标签通过自动生成方法(如使用Command-A模型和FastText语言识别)获得,虽然可能存在不完美之处,但人工评估表明它们为大规模分析提供了可靠信号。分析显示,在后续训练阶段,显性文化信号往往减少,导致现代LLM数据集中文化表征呈现长尾分布。数据集可用于通过宝藏标记方法改进下游文化基准性能。数据集特征包括:inputs(原始文本内容)、language(文本语言)、culture_tag(文化子维度标签)、geolocation_tag(地理位置信息)、domain_tag(数据所属领域)、task_intent_tag(样本的预期目的或任务类型)、dataset_source(原始数据集来源,如CulturaX、Dolci Instruct SFT等)和custom_id(唯一标识符)。
This dataset, titled The Culture Funnel: You Cant Align What isnt in the Data, contains approximately 5.6 million culturally tagged samples designed to help researchers study and mitigate the cultural data funnel in Large Language Model (LLM) pipelines. It uses a multidimensional tagging framework to identify cultural signals, domains, geographic locations, and task specialization across pretraining, fine-tuning, alignment, and reasoning datasets. Tags for cultural dimensions (such as knowledge, preference, dynamics, bias, general culture, and no culture), domain (e.g., math, humanities, conversation, social sciences), task intent, and geographic location are obtained via automatic methods like prompting the Command-A model and FastText language identification. Although these annotations are automatically generated and may be imperfect, human evaluation shows they provide reliable signals for large-scale analysis. The analysis reveals that explicit cultural signals often diminish throughout post-training stages, resulting in a long-tail distribution of cultural representation in modern LLM datasets. The dataset can be used to improve downstream cultural benchmark performance following the treasure marking approach. Features include: inputs (raw text content), language (text language), culture_tag (cultural subdimension tags), geolocation_tag (geographic region information), domain_tag (domain of the data), task_intent_tag (intended purpose or task type of the sample), dataset_source (original dataset source, such as CulturaX, Dolci Instruct SFT, etc.), and custom_id (a unique identifier for each sample).




