SCINLI
收藏资源简介:
SCINLI是一个专为科学文本自然语言推理设计的大型数据集,包含107,412个从NLP和计算语言学领域的学术论文中提取的句子对。该数据集捕捉了科学文本的正式性,并引入了CONTRASTING和REASONING两个新类别,以更好地反映科学文献中的推理关系。数据集的创建过程利用了远监督方法,通过链接短语来识别句子间的语义关系。SCINLI特别适合评估科学领域的自然语言理解模型,旨在解决现有NLI数据集未覆盖的科学文本推理问题。
SCINLI is a large-scale dataset specifically designed for natural language inference (NLI) over scientific texts, comprising 107,412 sentence pairs extracted from academic papers in the fields of natural language processing (NLP) and computational linguistics. This dataset captures the formal nature of scientific texts, and introduces two new categories, CONTRASTING and REASONING, to better reflect the inferential relationships present in scientific literature. The dataset was constructed using distant supervision methods, where linking phrases are utilized to identify semantic relationships between sentences. SCINLI is particularly suitable for evaluating natural language understanding models in the scientific domain, aiming to address the issue of scientific text inference that is not covered by existing NLI datasets.

- 1SciNLI: A Corpus for Natural Language Inference on Scientific Text伊利诺伊大学芝加哥分校计算机科学系 · 2022年



