LeWiDi-2025
收藏资源简介:
LeWiDi-2025数据集是一个用于训练和评估自然语言处理模型的资源,它包含了四个数据集,分别是:Conversational Sarcasm Corpus (CSC)、MultiPICo dataset (MP)、VariErr NLI dataset (VEN)和Paraphrase Detection dataset (Par)。这些数据集覆盖了自然语言处理的多个领域,包括讽刺检测、自然语言推理、释义检测等,并且包含了多个语言的标注数据。数据集采用了不同的标注方案,包括类别标签和序数标签,旨在帮助模型学习人类判断中的差异和变化。此外,LeWiDi-2025还引入了两种互补的评价范式:软标签预测和视角主义预测,以及新的评价指标,以更好地评估模型处理差异的能力。
The LeWiDi-2025 dataset is a curated resource for training and evaluating natural language processing (NLP) models. It encompasses four constituent datasets: the Conversational Sarcasm Corpus (CSC), MultiPICo dataset (MP), VariErr Natural Language Inference (NLI) dataset (VEN), and Paraphrase Detection dataset (Par). These datasets cover multiple NLP subfields, including sarcasm detection, natural language inference, paraphrase detection, and more, and feature annotated data across multiple languages. Adopting diverse annotation schemas including categorical labels and ordinal labels, the dataset is designed to assist models in learning the discrepancies and variations inherent in human judgments. Additionally, LeWiDi-2025 introduces two complementary evaluation paradigms: soft label prediction and perspectivist prediction, alongside novel evaluation metrics to better gauge models' capabilities in handling such variations.

- 1LeWiDi-2025 at NLPerspectives: The Third Edition of the Learning with Disagreements Shared TaskFondazione Bruno Kessler, LMU Munich &MCML, Universit`a Milano Bicocca, Universit`a di Torino, University of Gothenburg, Queen Mary University of London, Utrecht University · 2025年



