silence-suzuki/FIRE-Bench-unverified
收藏资源简介:
FIRE-Bench是一个包含153个研究任务的基准测试,这些任务是从58篇学术论文中自动生成的。每个任务都包含一个研究问题和原始论文使用的资源(如模型、数据集、预算和约束条件),并要求代理设计和运行自己的实验。数据集适用于文本生成和问答任务,支持英语语言环境。每个任务包含唯一标识符、论文类型、研究问题、完整提示、真实答案和任务配置等字段。数据集来源包括HuggingFace、外部链接、合成数据和未知来源。数据集目前处于未验证状态,由LLM管道端到端生成,未经人工审核。
FIRE-Bench is a benchmark of 153 research tasks auto-generated from 58 academic papers. Each task hands an agent a research question plus the resources the original paper used (models, datasets, budget, constraints) and asks it to design and run its own experiments. The dataset is suitable for text-generation and question-answering tasks in English. Each task includes fields such as a unique identifier, paper type, research question, full prompt, ground-truth answer, and task configuration. Dataset sources include HuggingFace, external URLs, synthetic data, and unknown sources. The dataset is currently unverified, generated end-to-end by an LLM pipeline without human review.



