BUT-FIT/OARelatedWorkMetaEval
收藏资源简介:
OARelatedWork元评估数据集是一个基于BUT-FIT/OARelatedWork构建的人工标注元评估数据集,用于衡量相关工作总结生成任务的自动指标与人类判断之间的相关性。数据集包含两个配置:duels配置(约120行)提供同一目标论文的两个生成相关工作总结的成对比较,包括偏好、相关性、忠实性和语言流畅性等维度的标注;statements配置(约408行)提供从生成或参考部分采样的单个原子陈述及其事实性标签(如True、False、True, but wrong citation、Unverifiable)。涉及的系统包括primera_go_all、gpt_4o_mini_go_all和人类参考。数据集还包括目标论文的详细信息(如标题、摘要、作者、引用文献等),并支持文本生成和摘要任务的研究。
OARelatedWork Meta-Evaluation is a human-annotated meta-evaluation dataset built on top of BUT-FIT/OARelatedWork, used to measure how well automatic metrics for related-work section generation correlate with human judgement. The dataset includes two configurations: duels (approx. 120 rows) provides pairwise comparisons of two generated related-work sections for the same target paper, with annotations on dimensions like preference, relevance, faithfulness, and language; statements (approx. 408 rows) provides individual atomic statements sampled from generated or reference sections, along with factuality labels (e.g., True, False, True, but wrong citation, Unverifiable). The systems compared include primera_go_all, gpt_4o_mini_go_all, and human reference. The dataset also contains detailed information about target papers (e.g., title, abstract, authors, references) and is intended for tasks like text generation and summarization.



