ResearchClawBench
收藏资源简介:
ResearchClawBench是一个用于评估AI编码代理能否独立进行科学研究的基准测试,从读取原始数据到生成可发表质量的研究报告,并将其结果与真实人类撰写的论文进行严格对比。该数据集包含40个真实科学任务,涵盖天文学、化学、地球科学等10个学科领域,每个任务都基于已发表论文的完整实验数据集。数据集采用两阶段评估流程:自主研究阶段(AI代理需独立分析数据、编写代码并生成报告)和基于参考的评估阶段(使用多模态LLM法官根据精细检查清单对报告进行评分)。每个任务包含原始数据集、参考材料、任务说明和评估检查清单。数据集结构清晰,包含任务信息、原始数据、相关工作和目标研究论文等目录。该基准测试支持多种AI代理,并提供实时流式UI观察代理的研究过程。
ResearchClawBench is a benchmark dataset for evaluating whether AI coding agents can conduct independent scientific research, covering the full pipeline from reading raw data to generating publishable-quality research reports, with its results rigorously compared against papers written by real human researchers. This dataset comprises 40 real-world scientific tasks across 10 academic disciplines including astronomy, chemistry, earth sciences and other related fields, with each task built upon the complete experimental dataset from a published academic paper. The dataset adopts a two-stage evaluation workflow: the autonomous research phase, where the AI agent is required to independently analyze data, write code and generate research reports, and the reference-based evaluation phase, where a multimodal LLM judge scores the generated reports based on a detailed inspection checklist. Each individual task includes raw datasets, reference materials, task instructions and an evaluation checklist. The dataset features a well-organized structure, with dedicated directories for task information, raw data, related work, target research papers and other relevant contents. This benchmark supports multiple AI agents, and provides a real-time streaming UI to observe the research process of the agents.





