遇见数据集

csebuetnlp/xnli_bn

收藏
Hugging Face2022-08-21 更新2024-03-04 收录
官方服务:

资源简介:

--- annotations_creators: - machine-generated language_creators: - found multilinguality: - monolingual size_categories: - 100K<n<1M source_datasets: - extended task_categories: - text-classification task_ids: - natural-language-inference language: - bn license: - cc-by-nc-sa-4.0 --- # Dataset Card for `xnli_bn` ## Table of Contents - [Dataset Card for `xnli_bn`](#dataset-card-for-xnli_bn) - [Table of Contents](#table-of-contents) - [Dataset Description](#dataset-description) - [Dataset Summary](#dataset-summary) - [Supported Tasks and Leaderboards](#supported-tasks-and-leaderboards) - [Languages](#languages) - [Usage](#usage) - [Dataset Structure](#dataset-structure) - [Data Instances](#data-instances) - [Data Fields](#data-fields) - [Data Splits](#data-splits) - [Dataset Creation](#dataset-creation) - [Curation Rationale](#curation-rationale) - [Source Data](#source-data) - [Initial Data Collection and Normalization](#initial-data-collection-and-normalization) - [Who are the source language producers?](#who-are-the-source-language-producers) - [Annotations](#annotations) - [Annotation process](#annotation-process) - [Who are the annotators?](#who-are-the-annotators) - [Personal and Sensitive Information](#personal-and-sensitive-information) - [Considerations for Using the Data](#considerations-for-using-the-data) - [Social Impact of Dataset](#social-impact-of-dataset) - [Discussion of Biases](#discussion-of-biases) - [Other Known Limitations](#other-known-limitations) - [Additional Information](#additional-information) - [Dataset Curators](#dataset-curators) - [Licensing Information](#licensing-information) - [Citation Information](#citation-information) - [Contributions](#contributions) ## Dataset Description - **Repository:** [https://github.com/csebuetnlp/banglabert](https://github.com/csebuetnlp/banglabert) - **Paper:** [**"BanglaBERT: Combating Embedding Barrier in Multilingual Models for Low-Resource Language Understanding"**](https://arxiv.org/abs/2101.00204) - **Point of Contact:** [Tahmid Hasan](mailto:tahmidhasan@cse.buet.ac.bd) ### Dataset Summary This is a Natural Language Inference (NLI) dataset for Bengali, curated using the subset of MNLI data used in XNLI and state-of-the-art English to Bengali translation model introduced **[here](https://aclanthology.org/2020.emnlp-main.207/).** ### Supported Tasks and Leaderboards [More information needed](https://github.com/csebuetnlp/banglabert) ### Languages * `Bengali` ### Usage ```python from datasets import load_dataset dataset = load_dataset("csebuetnlp/xnli_bn") ``` ## Dataset Structure ### Data Instances One example from the dataset is given below in JSON format. ``` { "sentence1": "আসলে, আমি এমনকি এই বিষয়ে চিন্তাও করিনি, কিন্তু আমি এত হতাশ হয়ে পড়েছিলাম যে, শেষ পর্যন্ত আমি আবার তার সঙ্গে কথা বলতে শুরু করেছিলাম", "sentence2": "আমি তার সাথে আবার কথা বলিনি।", "label": "contradiction" } ``` ### Data Fields The data fields are as follows: - `sentence1`: a `string` feature indicating the premise. - `sentence2`: a `string` feature indicating the hypothesis. - `label`: a classification label, where possible values are `contradiction` (0), `entailment` (1), `neutral` (2) . ### Data Splits | split |count | |----------|--------| |`train`| 381449 | |`validation`| 2419 | |`test`| 4895 | ## Dataset Creation The dataset curation procedure was the same as the [XNLI](https://aclanthology.org/D18-1269/) dataset: we translated the [MultiNLI](https://aclanthology.org/N18-1101/) training data using the English to Bangla translation model introduced [here](https://aclanthology.org/2020.emnlp-main.207/). Due to the possibility of incursions of error during automatic translation, we used the [Language-Agnostic BERT Sentence Embeddings (LaBSE)](https://arxiv.org/abs/2007.01852) of the translations and original sentences to compute their similarity. All sentences below a similarity threshold of 0.70 were discarded. ### Curation Rationale [More information needed](https://github.com/csebuetnlp/banglabert) ### Source Data [XNLI](https://aclanthology.org/D18-1269/) #### Initial Data Collection and Normalization [More information needed](https://github.com/csebuetnlp/banglabert) #### Who are the source language producers? [More information needed](https://github.com/csebuetnlp/banglabert) ### Annotations [More information needed](https://github.com/csebuetnlp/banglabert) #### Annotation process [More information needed](https://github.com/csebuetnlp/banglabert) #### Who are the annotators? [More information needed](https://github.com/csebuetnlp/banglabert) ### Personal and Sensitive Information [More information needed](https://github.com/csebuetnlp/banglabert) ## Considerations for Using the Data ### Social Impact of Dataset [More information needed](https://github.com/csebuetnlp/banglabert) ### Discussion of Biases [More information needed](https://github.com/csebuetnlp/banglabert) ### Other Known Limitations [More information needed](https://github.com/csebuetnlp/banglabert) ## Additional Information ### Dataset Curators [More information needed](https://github.com/csebuetnlp/banglabert) ### Licensing Information Contents of this repository are restricted to only non-commercial research purposes under the [Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License (CC BY-NC-SA 4.0)](https://creativecommons.org/licenses/by-nc-sa/4.0/). Copyright of the dataset contents belongs to the original copyright holders. ### Citation Information If you use the dataset, please cite the following paper: ``` @misc{bhattacharjee2021banglabert, title={BanglaBERT: Combating Embedding Barrier in Multilingual Models for Low-Resource Language Understanding}, author={Abhik Bhattacharjee and Tahmid Hasan and Kazi Samin and Md Saiful Islam and M. Sohel Rahman and Anindya Iqbal and Rifat Shahriyar}, year={2021}, eprint={2101.00204}, archivePrefix={arXiv}, primaryClass={cs.CL} } ``` ### Contributions Thanks to [@abhik1505040](https://github.com/abhik1505040) and [@Tahmid](https://github.com/Tahmid04) for adding this dataset.

提供机构:
csebuetnlp
原始信息汇总

数据集概述

数据集名称

xnli_bn

语言

  • Bengali

许可证

  • cc-by-nc-sa-4.0

数据集大小

  • 100K<n<1M

任务类别

  • text-classification

具体任务

  • natural-language-inference

数据集结构

  • 数据实例:包含sentence1, sentence2, label三个字段。

    • sentence1: 字符串,表示前提。
    • sentence2: 字符串,表示假设。
    • label: 分类标签,可能值为contradiction (0), entailment (1), neutral (2)。
  • 数据分割:

    • train: 381449条
    • validation: 2419条
    • test: 4895条

数据集创建

使用方法

python from datasets import load_dataset dataset = load_dataset("csebuetnlp/xnli_bn")

引用信息

@misc{bhattacharjee2021banglabert, title={BanglaBERT: Combating Embedding Barrier in Multilingual Models for Low-Resource Language Understanding}, author={Abhik Bhattacharjee and Tahmid Hasan and Kazi Samin and Md Saiful Islam and M. Sohel Rahman and Anindya Iqbal and Rifat Shahriyar}, year={2021}, eprint={2101.00204}, archivePrefix={arXiv}, primaryClass={cs.CL} }

搜集汇总
数据集介绍
csebuetnlp/xnli_bn 数据集图片
构建方式
在自然语言处理领域,跨语言理解任务常面临低资源语言的挑战,xnli_bn数据集的构建正是为了应对孟加拉语资源的稀缺性。该数据集基于XNLI框架,从MultiNLI训练数据中选取子集,通过先进的英译孟模型进行自动翻译。为确保翻译质量,研究团队采用语言无关的BERT句子嵌入技术,计算原文与译文之间的语义相似度,并设定0.70的阈值,剔除相似度较低的样本,从而构建出规模约38万条的高质量自然语言推理数据集。
特点
作为孟加拉语自然语言推理任务的重要资源,xnli_bn数据集展现出鲜明的语言学特征。其数据实例包含前提与假设两个文本字段,标注涵盖矛盾、蕴含和中性三类逻辑关系,全面覆盖自然语言推理的核心范畴。数据集规模适中,提供训练、验证和测试的标准划分,支持模型开发与评估的完整性。同时,数据以纯文本形式呈现,结构清晰,便于直接应用于预训练或微调任务,为低资源语言理解研究提供了可靠的基础。
使用方法
在低资源语言模型研究中,xnli_bn数据集为孟加拉语自然语言推理任务提供了标准化实验平台。用户可通过Hugging Face的datasets库直接加载数据,快速获取训练、验证和测试分割。数据字段包括sentence1、sentence2和label,可直接输入模型进行序列分类或句子对关系预测。研究者可利用该数据集评估跨语言模型的迁移性能,或作为预训练语料增强模型对孟加拉语的理解能力,推动低资源语言处理技术的发展。
背景与挑战
背景概述
在自然语言处理领域,跨语言理解一直是核心研究议题,尤其对于低资源语言如孟加拉语,其语义推理任务面临数据稀缺的困境。xnli_bn数据集应运而生,由孟加拉国工程技术大学计算机科学与工程系的研究团队于2021年创建,旨在扩展XNLI框架至孟加拉语,以支持自然语言推理任务。该数据集基于MultiNLI训练数据,通过先进的英译孟模型转化而成,其诞生显著提升了孟加拉语在预训练模型中的表征能力,为低资源语言的语义理解研究提供了关键基准。
当前挑战
xnli_bn数据集所针对的自然语言推理任务,在孟加拉语中面临语义细微差别捕捉的挑战,例如语言结构差异导致的逻辑关系歧义。在构建过程中,自动翻译机制可能引入语义失真,团队通过LaBSE嵌入相似度阈值过滤低质量样本,但翻译误差与文化特定表达的保留仍构成潜在局限。此外,数据规模虽达数十万,但源于单一翻译管道,可能限制模型对语言多样性的泛化能力。
常用场景
经典使用场景
在自然语言处理领域,孟加拉语作为低资源语言,其语义理解任务长期面临数据稀缺的挑战。xnli_bn数据集通过提供大规模、高质量的孟加拉语自然语言推理标注数据,成为该语言语义关系建模的基准资源。该数据集典型应用于训练和评估孟加拉语预训练语言模型,如BanglaBERT,通过判断前提句与假设句之间的蕴含、矛盾或中立关系,系统检验模型对语言逻辑的深层理解能力。
解决学术问题
该数据集有效缓解了低资源语言在自然语言推理任务上的数据匮乏困境,为跨语言语义理解研究提供了关键支撑。其构建过程中采用的翻译过滤机制,通过LaBSE嵌入相似度阈值控制,显著提升了翻译数据的语义保真度,从而解决了自动翻译引入噪声的学术难题。这一工作推动了低资源语言处理从依赖跨语言迁移向构建本土化高质量数据集的范式转变,对促进语言技术公平性具有深远意义。
衍生相关工作
围绕xnli_bn数据集衍生出一系列重要的研究工作,其中最具代表性的是BanglaBERT预训练模型的开发。该模型通过在该数据集上进行微调,显著提升了孟加拉语理解任务的性能指标。后续研究进一步拓展了数据集的利用方式,包括基于该数据集的跨语言对比学习框架构建、低资源语言推理能力评估体系建立等,这些工作共同推动了南亚语言处理研究社区的繁荣发展。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务