OpenMEVA
收藏资源简介:
OpenMEVA是由清华大学的研究团队开发的一个用于评估开放式故事生成度量的基准数据集。该数据集包含2000条经过人工标注的故事数据,这些数据来源于两个广泛使用的故事语料库:ROCStories和WritingPrompts。OpenMEVA旨在通过提供全面的测试套件来评估度量的能力,包括与人类判断的相关性、对不同模型输出和数据集的泛化能力、判断故事连贯性的能力以及对扰动的鲁棒性。数据集的创建过程涉及手动标注和自动构造测试示例,以确保数据的质量和多样性。OpenMEVA的应用领域主要集中在自然语言生成(NLG)模型的评估和改进,特别是在开放式故事生成任务中,旨在解决现有自动度量与人类评估之间相关性差的问题。
OpenMEVA is a benchmark dataset developed by the research team at Tsinghua University for evaluating open-ended story generation metrics. This dataset contains 2000 manually annotated story samples sourced from two widely adopted story corpora: ROCStories and WritingPrompts. OpenMEVA aims to evaluate the performance of metrics by providing a comprehensive test suite, covering correlations with human judgments, generalization ability across different model outputs and datasets, capability of assessing story coherence, and robustness against perturbations. The construction of this dataset involves manual annotation and automatic test example generation to ensure data quality and diversity. The application scenarios of OpenMEVA are mainly focused on the evaluation and improvement of natural language generation (NLG) models, particularly in the open-ended story generation task, aiming to address the issue of poor correlation between existing automatic metrics and human evaluations.

- 1OpenMEVA: A Benchmark for Evaluating Open-ended Story Generation Metrics清华大学 · 2021年



