遇见数据集

Datasets for "ESNLIR: A Spanish Multi-Genre Dataset with Causal Relationships"

收藏
Zenodo2025-03-13 更新2026-05-26 收录
官方服务:

资源简介:

ESNLIR: A Spanish Multi-Genre Dataset with Causal Relationships These are the datasets for the paper ESNLIR: A Spanish Multi-Genre Dataset with Causal Relationships. Dataset dictionary This repository contains the splits that resulted from the research project "ESNLIR: A Spanish Multi-Genre Dataset with Causal Relationships". All the splits are in JSONL format and have the same fields per example: sentence_1: First sentence of the pair. sentence_2: Second sentence of the pair. connector: Linking phrase used to extract pair. connector_type: NLI label, between "contrasting", "entailment", "reasoning" or "neutral" extraction_strategy: "linking_phrase" for "contrasting", "entailment", "reasoning" and "none" for neutral. distance: How many sentences before the connector is the sentence_1 sentence_1_position: Number of sentence for sentence_1 in the source document sentence_1_paragraph: Number of paragraph for sentence_1 in the source document sentence_2_position: Number of sentence for sentence_2 in the source document sentence_2_paragraph: Number of paragraph for sentence_2 in the source document id: Unique identifier for the example dataset: Source corpus of the pair. Metadata of corpus, including source can be found in dataset_metadata.xlsx. genre: Writing genre of the dataset. domain: Domain genre of the dataset. Example: {"sentence_1":"sefior Bcajavides no es moderado, tampoco lo convertirse e\u00f1 declarada divergencia de miras polileido en griego","sentence_2":"era mayor claricomentarios, as\u00ed de los peri\u00f3dicos como de los homes dado \u00e1 la voluntad de los hombres, sin que sobreticas","connector":"por consiguiente,","connector_type":"reasoning","extraction_strategy":"linking_phrase","distance":1.0,"sentence_1_paragraph":4,"sentence_1_position":86,"sentence_2_paragraph":4,"sentence_2_position":87,"id":"esnews__spanish_pd_news__531537","dataset":"esnews__spanish_pd_news","genre":"news","domain":"spanish_public_domain_news"} Dataset files ESNLIR_datasets.zip: Contains the splits used for BERT-based model training, validation and testing, including stress test splits. labeled_final_dataset.jsonl: Is the final test dataset with 974 examples selected by human majority label matching the original linking phrase label.

ESNLIR:带因果关系的西班牙语多体裁数据集 本数据集配套于论文《ESNLIR:带因果关系的西班牙语多体裁数据集》。 数据集字典说明 本仓库包含该研究项目「ESNLIR:带因果关系的西班牙语多体裁数据集」产出的数据集划分文件,所有划分文件均采用JSONL格式,每条样本包含如下字段: sentence_1:语句对中的第一句。 sentence_2:语句对中的第二句。 connector:用于提取该语句对的连接短语。 connector_type:自然语言推理(Natural Language Inference, NLI)标签,可选值为"contrasting"、"entailment"、"reasoning"或"neutral"。 extraction_strategy:提取策略,针对"contrasting"、"entailment"、"reasoning"三类标签为"linking_phrase"(连接短语法),针对"neutral"标签为"none"(无)。 distance:语句1与连接短语之间相隔的语句数。 sentence_1_position:语句1在源文档中的句子序号。 sentence_1_paragraph:语句1在源文档中的段落序号。 sentence_2_position:语句2在源文档中的句子序号。 sentence_2_paragraph:语句2在源文档中的段落序号。 id:该样本的唯一标识符。 dataset:该语句对的来源语料库。语料库元数据(包括来源信息)可参见dataset_metadata.xlsx文件。 genre:数据集的写作体裁。 domain:数据集所属领域。 样本示例: {"sentence_1":"sefior Bcajavides no es moderado, tampoco lo convertirse eñ declarada divergencia de miras polileido en griego","sentence_2":"era mayor claricomentarios, así de los periòdicos como de los homes dado á la voluntad de los hombres, sin que sobreticas","connector":"por consiguiente,","connector_type":"reasoning","extraction_strategy":"linking_phrase","distance":1.0,"sentence_1_paragraph":4,"sentence_1_position":86,"sentence_2_paragraph":4,"sentence_2_position":87,"id":"esnews__spanish_pd_news__531537","dataset":"esnews__spanish_pd_news","genre":"news","domain":"spanish_public_domain_news"} 数据集文件 ESNLIR_datasets.zip:包含用于基于BERT模型的训练、验证与测试的数据集划分文件,其中还包含压力测试划分集。 labeled_final_dataset.jsonl:为最终测试数据集,包含974条经人工多数投票标注、与原始连接短语标签一致的样本。

提供机构:
Arxiv
创建时间:
2025-03-10
二维码
社区交流群
二维码
科研交流群
商业服务