cultural-benchmark-annotations-final
收藏资源简介:
该数据集是一个用于系统性评估其他人工智能基准测试(benchmark)的元数据集。它并非包含原始任务数据,而是对一系列现有基准测试在多维度指标上的表现进行评估和量化后生成的评估结果集合。数据集的核心关注点在于分析基准测试在语言多样性、地域与文化表征、数据创建透明度、标注流程公平性以及文化偏见意识等方面的实践水平。每个数据条目对应一个被评估的基准测试,并包含以下主要维度的详细信息:1) 基准标识与引用信息;2) 基准所覆盖的语言、地理区域(大洲、国家、地区)、方言、文字以及是否包含服务不足群体;3) 基准中文化内容的深度与广度,涉及价值观、宗教、社会规范、叙事、流行文化等多个文化子主题;4) 数据创建过程的文档化程度、数据来源、筛选方法及质量控制措施;5) 标注人员(如果涉及)的招募、多样性、报酬公平性、指南可用性及标注质量控制;6) 基准对文化的定义深度、偏见意识及缓解措施、公平性报告等;7) 基于上述维度计算出的详细评分体系,包括原始分、标准化分数以及数十个细粒度指标的分项得分。该数据集旨在为研究社区提供一种标准化工具,用以衡量和比较不同基准测试在包容性、代表性和伦理实践方面的表现,促进人工智能评估向更公平、更多元化的方向发展。数据集共包含67个评估样本。
This dataset is a meta-dataset for systematically evaluating other artificial intelligence benchmarks. It does not contain original task data, but rather a collection of evaluation results generated by assessing and quantifying the performance of a series of existing benchmarks across multiple dimensions. The core focus of the dataset is to analyze the practical level of benchmarks in terms of language diversity, geographical and cultural representation, transparency in data creation, fairness in annotation processes, and awareness of cultural biases. Each data entry corresponds to an evaluated benchmark and includes detailed information in the following main dimensions: 1) Benchmark identification and citation information; 2) Languages covered by the benchmark, geographic regions (continents, countries, regions), dialects, scripts, and inclusion of underserved groups; 3) Depth and breadth of cultural content in the benchmark, involving cultural sub-themes such as values, religion, social norms, narratives, and popular culture; 4) Documentation level of the data creation process, data sources, filtering methods, and quality control measures; 5) Recruitment of annotators (if involved), diversity, fairness of compensation, availability of guidelines, and quality control of annotations; 6) Depth of cultural definition in the benchmark, bias awareness and mitigation measures, fairness reporting, etc.; 7) Detailed scoring system calculated based on the above dimensions, including raw scores, standardized scores, and sub-scores for dozens of fine-grained indicators. The dataset aims to provide the research community with a standardized tool to measure and compare the performance of different benchmarks in terms of inclusivity, representativeness, and ethical practices, promoting the development of more fair and diverse artificial intelligence evaluation. The dataset contains a total of 67 evaluation samples.
数据集概述:Cultural Benchmark Annotations Final
基本信息
- 数据集名称:Cultural Benchmark Annotations Final
- 数据集地址:https://huggingface.co/datasets/Josefine245/cultural-benchmark-annotations-final
- 数据集大小:下载大小为 154,671 字节,数据集大小为 169,613 字节
- 数据分割:仅包含训练集(
train),共 70 个样本
数据特征
数据集包含丰富的特征字段,主要分为以下几类:
-
基准与引用信息:
benchmark_id:基准 IDpaper_citation:论文引用timestamp_utc:UTC 时间戳
-
代表性维度:
- 语言、大洲、国家、方言、文字等(如
rep_lang_languages_list、rep_continents_list、rep_countries_list、rep_dialects、rep_scripts) - 代表性不足群体标志及列表(
rep_underserved_groups_flag、rep_underserved_groups_list) - 是否基于理论、理论来源、国家 vs 文化区分(
rep_theory_based、rep_theory_sources、rep_country_vs_culture)
- 语言、大洲、国家、方言、文字等(如
-
文化主题维度:
- 包含价值观、宗教、社会规范、叙事、流行文化、符号、仪式、服饰、饮食习惯、节日等 10 个文化主题的评分(如
cult_values、cult_religion等) - 文化主题数量、平衡性、合理性及来源(
cult_topics_num、cult_topics_balance_reflected、cult_topics_justified、cult_topics_sources)
- 包含价值观、宗教、社会规范、叙事、流行文化、符号、仪式、服饰、饮食习惯、节日等 10 个文化主题的评分(如
-
数据创建与处理:
- 问题类型(
data_question_types) - 文档处理过程、创建模式、创建方法(如
data_doc_process_documented、data_creation_mode、data_creation_methods) - 数据源来源、外部参考、过滤与清洗(如
data_source_origin_documented、data_selection_external_refs、data_filtering_cleaned) - 文化相关性过滤、质量控制、答案一致性检查(
cultural_relevance_filter、quality_control、answer_consistency_check)
- 问题类型(
-
标注者信息:
- 标注者参与数量、选择文档化、要求定义、文化相关性、多样性维度、招募渠道、支付公平性、指南可用性、质量检查方法等(如
annotators_involved、ann_requirements_defined、annotators_culturally_relevant、annotator_diversity_dimensions、ann_recruitment_channels等)
- 标注者参与数量、选择文档化、要求定义、文化相关性、多样性维度、招募渠道、支付公平性、指南可用性、质量检查方法等(如
-
文化定义与偏差:
- 文化定义的存在性、深度、来源、概念基础(如
culture_definition_exists、culture_definition_depth、culture_definition_with_sources、culture_conceptual_basis) - 偏差意识、缓解措施、公平性与包容性报告、透明度与局限性(
bias_awareness、bias_mitigation_measures、fairness_inclusivity_reporting、transparency_dataset_available、transparency_limitations_reflected)
- 文化定义的存在性、深度、来源、概念基础(如
-
评分与加权维度:
- 原始分数、最大分数、归一化分数(
raw_score、max_score、normalized_score) - 各维度加权得分及总分(如
weighted_total、dim1_weighted至dim6_weighted) - 语法库覆盖、语言解析、地点分化、大洲、国家、地区、方言、文字、代表性不足群体、理论、文化等子维度的点数(如
grambank_coverage、grambank_points、location_points、continent_points、country_points、region_points、dialect_points、script_points等)
- 原始分数、最大分数、归一化分数(
数据用途
该数据集用于对文化基准进行详细标注与评分,涵盖文化主题、数据创建、标注者信息、文化定义、偏差与公平性等多个维度,适合用于评估和提升自然语言处理模型在文化敏感性方面的表现。





