BioProBench
收藏资源简介:
BioProBench是一个针对生物实验协议理解和推理的大型综合性数据集,包含了从27,000个原始协议中提取的近556,000个高质量结构化实例。数据集涵盖了五个核心任务:协议问答、步骤排序、错误纠正、协议生成和协议推理,旨在对语言模型在生物文本处理方面的能力进行全面的评估。数据集来源于6个权威在线资源,覆盖了16个生物学子领域,具有广泛的领域覆盖性和代表性。数据集的构建经历了数据收集、处理、任务实例生成和多层质量控制等多个阶段,确保了数据的准确性和可靠性。
BioProBench is a large-scale comprehensive dataset focused on biological experimental protocol understanding and reasoning, comprising nearly 556,000 high-quality structured instances extracted from 27,000 original protocols. It covers five core tasks: protocol question answering, step ordering, error correction, protocol generation and protocol reasoning, aiming to comprehensively evaluate the capabilities of language models in biological text processing. The dataset is sourced from six authoritative online resources, spans 16 subfields of biology, and features extensive domain coverage and representativeness. The construction of the dataset has gone through multiple stages including data collection, preprocessing, task instance generation and multi-tiered quality control, ensuring the accuracy and reliability of the dataset.

- 1BioProBench: Comprehensive Dataset and Benchmark in Biological Protocol Understanding and Reasoning北京大学 · 2025年



