遇见数据集

QizhiPei/BioMatrix-SFT

收藏
Hugging Face2026-05-06 更新2026-05-31 收录
官方服务:

资源简介:

BioMatrix-SFT是一个用于训练BioMatrix多模态基础模型的监督微调/指令调优语料库。该数据集覆盖了6个类别、80个下游任务,涵盖分子、蛋白质和相互作用三个生物实体范围,每个范围包括1D序列和3D结构模态。数据组织为7个配置,包括分子1D(SMILES和SELFIES)、分子3D、蛋白质1D、蛋白质3D、相互作用1D和相互作用3D。数据集采用统一的指令调优模式,每个示例包含instruction、input、output、history和source字段。生物分子内容通过BioMatrix的统一标记化方案序列化,并使用模态特定的控制标记包装,以支持多模态任务。数据集旨在用于复现或扩展BioMatrix的指令调优,以及基准测试其他多模态生物基础模型。

BioMatrix-SFT is a supervised fine-tuning/instruction-tuning corpus used to train BioMatrix, a multimodal foundation model. The dataset covers 80 downstream tasks across 6 categories, spanning three biological entity scopes (molecule, protein, interaction) and, within each scope, both 1D sequence and 3D structure modalities. It is organized into 7 configs, including molecule 1D (SMILES and SELFIES), molecule 3D, protein 1D, protein 3D, interaction 1D, and interaction 3D. The dataset follows a unified instruction-tuning schema, with each example containing fields for instruction, input, output, history, and source. Biomolecular content is serialized using BioMatrixs unified tokenization scheme and wrapped with modality-specific control tokens to support multimodal tasks. The dataset is intended for reproducing or extending BioMatrixs instruction tuning and benchmarking other multimodal biological foundation models.

提供机构:
QizhiPei
二维码
社区交流群
二维码
科研交流群
商业服务