EdisonScientific/bixbench_hypothesis
收藏资源简介:
BixBench-Hypothesis (BBH) 是原始 BixBench 生物信息学基准测试的一个衍生版本,旨在衡量原始 BixBench 未捕获的能力:在模糊性下的分析判断或追求开放式目标假设的能力,这些更接近科学家的实际工作方式。BBH 适应了原始胶囊框架,将每个任务重新配置为(数据、假设、协议)元组,并配以评估标准。协议是逐步指南,用于对数据进行分析以解决假设;标准是结构化预期输出集,用于评分代理的分析,包括1或2点的子任务和与正确支持或拒绝假设相关的5点最终目标。数据集包含51个假设驱动任务,其中30个主要使用Python,21个主要使用R。
BixBench-Hypothesis (BBH) is a derivation of the original BixBench benchmark for bioinformatics, intended to measure abilities not captured by BixBench: analytical judgment under ambiguity or the ability to pursue a hypothesis with an open-ended goal, which are closer to how scientists actually work. BBH adapts the original capsule framework, reconfiguring each task as a (data, hypothesis, protocol) tuple paired with an evaluation rubric. The protocol is a step-by-step guide for carrying out an analysis on the data to address the hypothesis. The rubric is a structured set of expected outputs used to score an agents analysis, consisting of 1- or 2-point subtasks plus a 5-point final objective tied to correctly supporting or rejecting the hypothesis. There are 51 hypothesis-driven tasks comprising BixBench-Hypothesis. 30 capsules have a primary language of Python, while 21 are primarily R.





