RECOM (Reddit Evaluation for Correspondence of Models)
收藏资源简介:
RECOM数据集是由斯坦福大学等机构联合构建的用于评估大语言模型在开放式问答任务中的基准数据集,其全称为Reddit Evaluation for Correspondence of Models。该数据集包含15,000条源自r/AskReddit子论坛的英文问题及对应的真实社区回复,数据规模总计约647,973条对比样本,所有数据均采集于2025年9月以确保免受训练数据污染。数据集通过Pushshift API收集原始帖文,并筛选高互动内容构建而成,旨在为评估模型生成内容与社区共识的匹配度提供纯净测试环境。该数据集主要应用于自然语言处理领域,用于探究自动评估指标在开放式、观点驱动型问答任务中的有效性-判别力权衡问题,为衡量语言模型与人类社区观点的一致性提供标准化的评估框架。
The RECOM dataset, whose full name is Reddit Evaluation for Correspondence of Models, is a benchmark dataset jointly constructed by Stanford University and other institutions for evaluating large language models (LLMs) in open-ended question answering tasks. It contains 15,000 English questions sourced from the r/AskReddit subreddit and their corresponding authentic community responses, with a total of approximately 647,973 paired comparison samples. All data was collected in September 2025 to avoid contamination from model training corpora. The dataset was built by collecting original posts via the Pushshift API and filtering for high-engagement content, aiming to provide a clean test environment for assessing the alignment between model-generated content and community consensus. Primarily applied in the field of natural language processing (NLP), this dataset is used to explore the effectiveness-discrimination trade-off of automatic evaluation metrics in open-ended, opinion-driven question answering tasks, and to provide a standardized evaluation framework for measuring the consistency between language models and human community viewpoints.




