Biology-Instructions
收藏资源简介:
Biology-Instructions是由上海人工智能实验室创建的首个大规模多组学生物序列指令调优数据集,涵盖DNA、RNA、蛋白质及多分子预测任务,共包含21个子任务。该数据集拥有超过300万条训练样本,旨在通过多样化的生物序列任务提升大语言模型的推理能力和对话流畅性。数据集的构建过程包括从高质量文献和竞赛中收集任务数据,并通过人工和AI生成问答模板,确保语言风格和语法多样性。该数据集的应用领域主要集中在生物序列分析,旨在解决大语言模型在生物序列理解任务中的性能瓶颈,推动多组学序列分析与大语言模型的深度融合。
Biology-Instructions is the first large-scale multi-omic biological sequence instruction tuning dataset developed by the Shanghai AI Laboratory. It covers DNA, RNA, protein, and multi-molecule prediction tasks, totaling 21 subtasks. This dataset contains over 3 million training samples, aiming to enhance the reasoning capabilities and conversational fluency of large language models (LLMs) through diverse biological sequence tasks. The dataset construction process includes collecting task data from high-quality scholarly literature and competitions, and generating question-answer templates via both manual creation and AI-assisted generation to ensure diversity in linguistic styles and grammatical structures. The main application fields of this dataset focus on biological sequence analysis, aiming to address the performance bottlenecks of large language models in biological sequence understanding tasks and promote the in-depth integration of multi-omic sequence analysis and large language models.




