PARAPHRASUS
收藏资源简介:
PARAPHRASUS是由苏黎世大学创建的一个多维度评估基准,旨在测试和选择释义检测模型。该数据集包含43976条数据,涵盖了不同语义和词汇相似度的句子对,用于评估模型在不同释义类型上的表现。数据集的创建过程包括从现有数据集中重新利用数据,以及创建两个新的数据集,其中一个是由专家标注的具有挑战性的非对抗性测试集。PARAPHRASUS的应用领域广泛,旨在解决释义检测模型在不同语境下的泛化能力和性能评估问题。
PARAPHRASUS is a multidimensional evaluation benchmark developed by the University of Zurich for testing and selecting paraphrase detection models. This dataset comprises 43,976 data samples, covering sentence pairs with varying degrees of semantic and lexical similarity, which are used to assess model performance across different paraphrase types. The construction of PARAPHRASUS involves repurposing data from existing datasets and creating two new datasets, one of which is a challenging non-adversarial test set annotated by domain experts. With wide-ranging application domains, this benchmark aims to address the challenges in evaluating the generalization ability and performance of paraphrase detection models across diverse contextual scenarios.
数据集概述
数据集列表
-
PAWS-X
Link: PAWS-X Dataset -
SICK-R
Link: SICK-R Dataset -
MSRPC
Link: Microsoft Research Paraphrase Corpus -
XNLI
Link: XNLI Dataset -
ANLI
Link: Adversarial NLI (ANLI) -
Stanford NLI (SNLI)
Link: SNLI Dataset -
STS Benchmark
Link: STS Benchmark -
OneStopEnglish Corpus
Link: OneStopEnglish Corpus
新增数据集
-
AMR Paraphrases
Link: AMR Paraphrases -
STS Benchmark (STS-H) with Human Annotation - Consensus (Column)
Link: STS Benchmark
许可证
该仓库继承自原始发布的许可证,所有使用的数据集均为公开可用。

- 1PARAPHRASUS : A Comprehensive Benchmark for Evaluating Paraphrase Detection Models苏黎世大学 · 2024年



