遇见数据集

labbench2

收藏
魔搭社区2026-04-28 更新2026-07-19 收录
官方服务:

资源简介:

[![arXiv](https://img.shields.io/badge/arXiv-2501.XXXXX-b31b1b.svg)](https://arxiv.org/abs/2501.XXXXX) # LABBench2 ```LABBench2``` is a benchmark for measuring real-world capabilities of AI systems performing scientific research tasks. It is an evolution of the [Language Agent Biology Benchmark (LAB-Bench)](https://arxiv.org/abs/2407.10362), comprising nearly 1,900 tasks that measure similar capabilities but in more realistic contexts. ```LABBench2``` provides a meaningful jump in difficulty over LAB-Bench (model-specific accuracy differences range from −26% to −46% across subtasks), underscoring continued room for improvement. LABBench2 aims to be a standard benchmark for evaluating and advancing AI capabilities in scientific research. **This repository** contains the dataset of benchmark tasks. We also provide a public evaluation harness for running any model or agent against the benchmark, which is available on [GitHub](https://github.com/EdisonScientific/labbench2). --- ## Changelog Notable changes to ```LABBench2``` will be documented here. We expect to update the datset only in the case of clear issues, and do not intend to meangingfully change the benchmark over time. **2026-03-13** - We corrected an inadvertent data issue with `sourcequality` tasks. This has resulted in an entirely new set of 150 tasks being incorporated into the dataset. Published results have been updated accordingly.

提供机构:
maas
创建时间:
2026-02-06
二维码
社区交流群
二维码
科研交流群
商业服务