DATASETRESEARCH
收藏资源简介:
DATASETRESEARCH是一个全面评估人工智能代理在按需数据集发现和综合方面的能力的基准。该基准包含了来自Huggingface和PaperswithCode的208个真实世界的数据集需求,涵盖了六大自然语言处理任务。数据集的构建过程首先从超过100万个候选数据集中筛选出208个实例,然后利用OpenAI的o3模型处理相关的README文件和数据样本,生成六维度的元数据。最后,o3模型合成这些元数据以生成对应的数据集需求。DATASETRESEARCH旨在评估搜索代理和推理代理在数据集发现和综合方面的能力,通过元数据评估、少样本性能评估和监督微调效果等三个评估方法来衡量代理系统的性能。
DATASETRESEARCH is a benchmark for comprehensively evaluating the capabilities of AI Agents in on-demand dataset discovery and synthesis. This benchmark includes 208 real-world dataset requirements sourced from Hugging Face and Papers with Code, covering six natural language processing tasks. The construction process of the benchmark first selects 208 instances from over one million candidate datasets, then utilizes OpenAI's o3 model to process the corresponding README files and data samples, generating six-dimensional metadata. Finally, the o3 model synthesizes these metadata to generate the corresponding dataset requirements. DATASETRESEARCH aims to evaluate the capabilities of search agents and reasoning agents in dataset discovery and synthesis, measuring the performance of agent systems through three evaluation methods: metadata evaluation, few-shot performance evaluation, and supervised fine-tuning effect.
- 1DatasetResearch: Benchmarking Agent Systems for Demand-Driven Dataset Discovery上海交通大学, SII, GAIR · 2025年



