labbench2
收藏资源简介:
[](https://arxiv.org/abs/2501.XXXXX) # LABBench2 ```LABBench2``` is a benchmark for measuring real-world capabilities of AI systems performing scientific research tasks. It is an evolution of the [Language Agent Biology Benchmark (LAB-Bench)](https://arxiv.org/abs/2407.10362), comprising nearly 1,900 tasks that measure similar capabilities but in more realistic contexts. ```LABBench2``` provides a meaningful jump in difficulty over LAB-Bench (model-specific accuracy differences range from −26% to −46% across subtasks), underscoring continued room for improvement. LABBench2 aims to be a standard benchmark for evaluating and advancing AI capabilities in scientific research. **This repository** contains the dataset of benchmark tasks. We also provide a public evaluation harness for running any model or agent against the benchmark, which is available on [GitHub](https://github.com/EdisonScientific/labbench2). --- ## Changelog Notable changes to ```LABBench2``` will be documented here. We expect to update the datset only in the case of clear issues, and do not intend to meangingfully change the benchmark over time. **2026-03-13** - We corrected an inadvertent data issue with `sourcequality` tasks. This has resulted in an entirely new set of 150 tasks being incorporated into the dataset. Published results have been updated accordingly.



