AutoScholarQuery, RealScholarQuery
收藏资源简介:
AutoScholarQuery是由字节跳动研究院创建的高质量合成数据集,专为AI领域的学术搜索任务设计。该数据集包含35,511条细粒度学术查询及其对应的论文,数据来源于ICLR、ICML、NeurIPS、ACL和CVPR等顶级AI会议的论文。数据集通过GPT-4生成学术查询,并仅保留可在arXiv上检索到的论文。AutoScholarQuery旨在通过强化学习优化PaSa模型,提升其在复杂学术查询中的表现。RealScholarQuery则是一个包含50条真实世界学术查询的基准数据集,用于评估PaSa在现实场景中的性能。该数据集通过人工收集和标注相关论文,确保查询与论文的相关性。两个数据集的应用领域主要集中在学术文献检索,旨在解决复杂学术查询的自动化处理问题。
AutoScholarQuery is a high-quality synthetic dataset developed by ByteDance Research, specifically designed for academic search tasks in the AI field. This dataset comprises 35,511 fine-grained academic queries and their corresponding papers, sourced from top-tier AI conference proceedings including ICLR, ICML, NeurIPS, ACL, and CVPR. The academic queries are generated via GPT-4, and only papers retrievable on arXiv are retained in the dataset. AutoScholarQuery is intended to optimize the PaSa model through reinforcement learning, thereby enhancing its performance on complex academic search queries. RealScholarQuery, on the other hand, is a benchmark dataset containing 50 real-world academic queries, used to evaluate the performance of PaSa in real-world scenarios. Relevant papers for this dataset are manually collected and annotated to ensure the relevance between each query and its matched papers. The application scope of both datasets primarily centers on academic literature retrieval, with the goal of addressing challenges in the automated processing of complex academic queries.

- 1PaSa: An LLM Agent for Comprehensive Academic Paper Search字节跳动研究院, 北京大学 · 2025年



