aspectsim/AspectSim-Evaluation-Benchmark
收藏资源简介:
AspectSim是一个大规模基于方面的文档对相似性评估基准。每个实例包含两篇完整文档、一个基于比较的自然语言方面以及一个人类可解释的相似性标签。该基准涵盖五个不同领域:新闻、观点、酒店评论、医学文献和科学同行评审,支持在多领域条件下评估基于方面的相似性模型。数据集包含约26,000个实例,使用GPT-4o作为可扩展的标注工具进行整理,并通过严格的人工标注验证,标签准确率达到94.2%(95%置信区间:93.3–95.1%)。
AspectSim is a large-scale aspect-conditioned document-pair similarity evaluation benchmark. Each instance consists of two full documents, a natural-language aspect on which the comparison is based, and a human-interpretable similarity label on an ordinal scale. The benchmark spans five diverse domains: news, opinion, hotel reviews, medical literature, and scientific peer reviews, enabling evaluation of aspect-aware similarity models under realistic multi-domain conditions. The dataset comprises approximately 26,000 instances curated using GPT-4o as a scalable annotation tool, with similarity labels validated through rigorous human annotation, achieving 94.2% label accuracy (95% CI: 93.3–95.1%).




