alee_datasets
收藏资源简介:
该数据集是一个多语言测试集,包含三个独立配置:alee_bq275、alee_f200和alee_mt61,每个配置针对不同的语言覆盖范围和任务设计。数据集核心特征是多语言文本字段,采用“语言代码_文字代码”(例如eng_Latn、cmn_Hans)或“语言代码_国家代码”(例如en_EN、ar_EG)的命名约定,覆盖从广泛使用语言到低资源语言的数百种语言变体。具体配置包括:alee_bq275包含275种语言/方言的字段,有864个样本;alee_f200包含200种语言/方言的字段,有829个样本;alee_mt61包含61种语言/方言的字段,样本数未明确。每个样本除了多语言文本字段外,还包含元数据字段如id、domain、register、tags、level、split等。特别地,每个配置都包含一组以“negative”为后缀的英语字段(如eng_RoleSwap_negative),表明数据集可能用于涉及文本转换、对比或对抗性示例的任务。所有配置仅包含测试集(test split),适用于多语言自然语言处理任务的评估,例如机器翻译质量评估、跨语言文本分类、多语言语义相似度计算或对抗性文本生成研究。
This dataset is a multilingual test set comprising three independent configurations: alee_bq275, alee_f200, and alee_mt61, each designed for different language coverage ranges and task specifications. The core feature of the dataset is multilingual text fields, following naming conventions such as language code_script code (e.g., eng_Latn, cmn_Hans) or language code_country code (e.g., en_EN, ar_EG), covering hundreds of language variants from widely used languages to low-resource languages. Specific configurations include: alee_bq275 contains fields for 275 languages/dialects with 864 samples; alee_f200 contains fields for 200 languages/dialects with 829 samples; alee_mt61 contains fields for 61 languages/dialects, with the sample count not explicitly stated. Each sample includes multilingual text fields along with metadata fields such as id, domain, register, tags, level, split, etc. Notably, each configuration includes a set of English fields with a negative suffix (e.g., eng_RoleSwap_negative), indicating potential use in tasks involving text transformation, contrast, or adversarial examples. All configurations consist solely of a test split and are suitable for evaluating multilingual natural language processing tasks, such as machine translation quality assessment, cross-lingual text classification, multilingual semantic similarity computation, or adversarial text generation research.
数据集概述:Andrianos/alee_datasets
该数据集集包含三个不同的配置(config),每个配置包含一个测试集(test split),用于评估多语言或对抗性样本下的模型表现。
1. 配置:alee_bq275
- 描述:包含一个测试集,共 864 个样本,数据总大小为 43.1 MB(下载大小约 23.3 MB)。
- 特征:包含 5 个英文对抗性特征(如
eng_Latn、eng_RoleSwap_negative等)、大量不同语言字段(约 250 种语言变体,如aar_Latn、zul_Latn等),以及元数据字段(id、uniq_id、domain、register、tags、level、split、par_id、par_comment、orig_text、newline_next)。所有非元数据字段均为字符串类型。 - 语言覆盖:涵盖众多语言,包括但不限于英语、阿拉伯语、中文(简体/繁体)、印地语、西班牙语、法语、俄语、日语等,且多种语言包含不同的文字系统(如拉丁、阿拉伯、天城文、西里尔文等)。
2. 配置:alee_f200
- 描述:包含一个测试集,共 829 个样本,数据总大小为 32.1 MB(下载大小约 17.6 MB)。
- 特征:包含 5 个英文对抗性特征(如
eng_Latn、eng_RoleSwap_negative等)、众多语言字段(约 200 种语言变体,如ace_Arab、zul_Latn等),以及元数据字段(id、URL、domain、topic、has_image、has_hyperlink、SIB_CATEGORY)。所有非元数据字段均为字符串类型。 - 语言覆盖:涵盖约 200 种语言变体,包括非洲、亚洲、欧洲、美洲等地区的语言,如斯瓦希里语、祖鲁语、亚美尼亚语、格鲁吉亚语、缅甸语等。
3. 配置:alee_mt61
- 描述:仅提供特征定义,未指定具体样本数量或大小。包含一个测试集(未提供样本数和大小),但根据结构推测其规模与其他配置类似。
- 特征:包含 5 个英文对抗性特征(
en_EN、en_RoleSwap_negative、en_PolarityNegation_negative、en_AntonymRepl_negative、en_HypernymSub_negative),以及 61 种语言变体(如ar_EG、fr_FR、zh_CN、hi_IN等),均为字符串类型。特征名称采用更紧凑的语言-地区格式(如ar_EG表示埃及阿拉伯语)。 - 语言覆盖:覆盖 61 种语言-地区组合,包括阿拉伯语(埃及、沙特)、英语、西班牙语、法语(加拿大、法国)、中文、印地语、日语、韩语、俄语、葡萄牙语等常见语言,以及一些低资源语言。
数据集通用特点
- 测试集专用:所有配置仅提供测试集(test split),无训练或验证集划分。
- 对抗性评估:每个配置都包含 5 个英文对抗性特征(
RoleSwap_negative、PolarityNegation_negative、AntonymRepl_negative、HypernymSub_negative),可能用于评估模型在角色交换、极性否定、反义词替换和上位词替换下的鲁棒性。 - 用途:该数据集主要面向多语言和对抗性场景下的模型评估,尤其是机器翻译、跨语言理解或分类任务。




