dnagpt/biopaws-2
收藏资源简介:
BioPAWS-2是首个用于生物基础模型的聊天形式指令调优数据集和基准测试,它将整个生物序列分析领域(包括分类、回归、检索、结构、变异效应、跨模态、推理和多模态任务)重新表达为一个统一的聊天形式指令调优语料库。该数据集同时作为训练资源(SFT语料库)和基准测试:任何模型(如具有自定义头的专用蛋白质/DNA语言模型、零样本回答的通用LLM,或在此语料库上微调的LLM)都可以在一个共同轴上进行训练和评估。数据集包含约306K个示例,覆盖22个任务和9个任务家族,采用统一的聊天格式,设计为可扩展,并支持双重评估协议(零样本QA和微调后评估)。任务家族包括配对同源性、蛋白质功能、DNA分类、变异效应、结构分析、跨模态、推理链、生物医学QA和多模态任务。每个示例以JSON聊天记录形式存储,包含指令、候选标签和答案,适用于分类和回归任务。数据集整合了多个来源的数据,如BioPAWS、LLaMA-Gene、ProteinGym、UniProt、BixBench等,每个记录带有自己的许可和来源字段。
BioPAWS-2 is the first chat-form instruction-tuning dataset and benchmark for biological foundation models, re-expressing the entire landscape of biological sequence analysis — including classification, regression, retrieval, structure, variant effect, cross-modal, reasoning, and multimodal tasks — as a single chat-form instruction-tuning corpus. It serves simultaneously as a training resource (SFT corpus) and a benchmark: any model, such as a specialized protein/DNA language model with a custom head, a general-purpose LLM answering zero-shot, or an LLM fine-tuned on this corpus, can be trained and evaluated on one common axis. The dataset contains approximately 306K examples across 22 tasks and 9 task families, using a uniform chat format and designed for extensibility, with a dual evaluation protocol (zero-shot QA and fine-tune-then-evaluate). Task families cover pairwise homology, protein function, DNA classification, variant effect, structure analysis, cross-modal tasks, reasoning chains, biomedical QA, and multimodal tasks. Each example is stored as a JSON chat record with instructions, candidate choices, and answers, suitable for classification and regression tasks. It consolidates data from multiple sources like BioPAWS, LLaMA-Gene, ProteinGym, UniProt, BixBench, etc., with each record carrying its own license and source fields.




