遇见数据集

alvations/xnli-15way

收藏
Hugging Face2023-05-10 更新2024-03-04 收录
官方服务:

资源简介:

XNLI consists of 10k English sentences translated into 14 languages: ar: Arabic bg: Bulgarian de: German el: Greek es: Spanish fr: French hi: Hindi ru: Russian sw: Swahili th: Thai tr: Turkish ur: Urdu vi: Vietnamese zh: Chinese (Simplified) The XNLI 15-way parallel corpus can be used for Machine Translation as evaluation sets, in particular for low-resource languages such as Swahili or Urdu. We provide two files: xnli.15way.orig.tsv and xnli.15way.tok.tsv containing respectively the original and the tokenized version of the corpus. The files consist of 15 tab-separated columns, each corresponding to one language as indicated by the header. Please consider citing the following paper if using this dataset: @InProceedings{conneau2018xnli, author = "Conneau, Alexis and Rinott, Ruty and Lample, Guillaume and Williams, Adina and Bowman, Samuel R. and Schwenk, Holger and Stoyanov, Veselin", title = "XNLI: Evaluating Cross-lingual Sentence Representations", booktitle = "Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing", year = "2018", publisher = "Association for Computational Linguistics", location = "Brussels, Belgium", }

XNLI 包含10000条英语句子,这些句子已被翻译为14种语言: ar:阿拉伯语(Arabic) bg:保加利亚语(Bulgarian) de:德语(German) el:希腊语(Greek) es:西班牙语(Spanish) fr:法语(French) hi:印地语(Hindi) ru:俄语(Russian) sw:斯瓦希里语(Swahili) th:泰语(Thai) tr:土耳其语(Turkish) ur:乌尔都语(Urdu) vi:越南语(Vietnamese) zh:简体中文(Chinese (Simplified)) XNLI 15路平行语料库可作为机器翻译的评估集使用,尤其适用于斯瓦希里语、乌尔都语等低资源语言。 我们提供两个文件:xnli.15way.orig.tsv 与 xnli.15way.tok.tsv,二者分别收录该语料库的原始版本与Token化版本。这些文件包含15个制表符分隔的列,每一列对应一种语言,与表头标注的信息一致。 若使用该数据集,请引用以下论文: @InProceedings{conneau2018xnli, author = "Conneau, Alexis and Rinott, Ruty and Lample, Guillaume and Williams, Adina and Bowman, Samuel R. and Schwenk, Holger and Stoyanov, Veselin", title = "XNLI: 评估跨语言句子表示", booktitle = "2018年自然语言处理经验方法会议论文集", year = "2018", publisher = "国际计算语言学协会(Association for Computational Linguistics)", location = "比利时布鲁塞尔", }

提供机构:
alvations
原始信息汇总

数据集概述

数据集名称

XNLI

数据集内容

  • 语言种类:包含15种语言,具体包括:

    • Arabic (ar)
    • Bulgarian (bg)
    • German (de)
    • Greek (el)
    • Spanish (es)
    • French (fr)
    • Hindi (hi)
    • Russian (ru)
    • Swahili (sw)
    • Thai (th)
    • Turkish (tr)
    • Urdu (ur)
    • Vietnamese (vi)
    • Chinese (Simplified) (zh)
  • 数据用途:适用于机器翻译评估,特别是针对低资源语言如Swahili或Urdu。

数据集文件

  • 文件1xnli.15way.orig.tsv - 包含原始语料。
  • 文件2xnli.15way.tok.tsv - 包含已分词的语料。
  • 文件格式:15列,每列对应一种语言,以制表符分隔。

引用信息

  • 论文:Conneau, Alexis et al. "XNLI: Evaluating Cross-lingual Sentence Representations". Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, 2018.
  • 出版信息:Brussels, Belgium.
二维码
社区交流群
二维码
科研交流群
商业服务