apodex/Deep-Research-Benchmarks
收藏资源简介:
Deep Research Benchmarks是一个用于深度研究评估的基准数据集集合,主要用于在AgentHarness框架中以标准ReAct模式评估Apodex-1.0模型。该数据集包含多个公开的深度研究基准,例如BrowseComp、BrowseComp-ZH、XBench-DeepResearch-202510、DeepSearchQA、WideSearch、FrontierScience-Research、FrontierScience-Olympiad和SUPERChem-Text。其中,FrontierScience-Research和FrontierScience-Olympiad由Apodex团队构建。数据集以密码保护的形式分发,以防止网络爬虫和训练数据管道索引其答案内容,这旨在避免数据污染而非安全控制。数据集不包括Humanitys Last Exam(HLE),因为其许可证禁止重新分发。
Deep Research Benchmarks is a collection of public deep-research benchmarks used to evaluate the Apodex-1.0 model in standard ReAct mode within the AgentHarness framework. The dataset includes multiple benchmarks such as BrowseComp, BrowseComp-ZH, XBench-DeepResearch-202510, DeepSearchQA, WideSearch, FrontierScience-Research, FrontierScience-Olympiad, and SUPERChem-Text, with FrontierScience-Research and FrontierScience-Olympiad constructed by the Apodex team. It is distributed as a password-protected bundle to prevent web crawlers and training-data pipelines from indexing the answers, serving as contamination protection rather than a security measure. The dataset does not include Humanitys Last Exam (HLE) due to licensing restrictions that prohibit redistribution.





