dnagpt/omnigene4-sft-data
收藏资源简介:
OmniGene-4 SFT语料库是一个用于OmniGene-4和OmniGene-4-MM家族的监督微调数据集,旨在支持生物信息学领域的文本生成和问答任务。该数据集包含多个文件,总行数约在100K到1M之间,覆盖八个任务家族:蛋白质同源性(BioPAWS)、DNA、结构(3Di/DSSP)、细胞生物学、分子、突变、结构预测和一般生物问答。数据以JSONL格式存储,每行包含instruction、input、output和category等字段,部分行还包括task_name和subtask。数据集用于训练和评估生物语言模型,如Bio-SFT v2到v5版本,以及OmniGene-4-MM的阶段2/3 LoRA训练。文件包括训练集、评估集、种子数据和子集,如细胞生物学和分子SFT子集。数据集基于CC-BY-4.0许可证,支持英文和中文。
The OmniGene-4 SFT Corpus is a supervised fine-tuning dataset tailored for the OmniGene-4 and OmniGene-4-MM model families, intended to support text generation and question answering tasks within the bioinformatics domain. Comprising multiple files with a total line count ranging from approximately 100K to 1M, the dataset covers eight task families: Protein Homology (BioPAWS), DNA, Structure (3Di/DSSP), Cell Biology, Molecular Biology, Mutation, Structure Prediction, and General Bio-QA. All data is stored in JSONL format, where each line contains fields including instruction, input, output, and category; a subset of lines additionally includes task_name and subtask. This dataset is utilized for training and evaluating bio-language models such as Bio-SFT v2 through v5, as well as Phase 2/3 LoRA training for OmniGene-4-MM. The included files cover training splits, validation splits, seed data, and subsets like the Cell Biology and Molecular SFT subsets. The dataset is released under the CC-BY-4.0 license and supports both English and Chinese languages.




