Mol-Instructions
收藏资源简介:
Mol-Instructions是由浙江大学计算机科学与技术学院创建的一个大规模生物分子指令数据集,旨在通过分子导向指令、蛋白质导向指令和生物分子文本指令三个核心组件,提高大型语言模型在生物分子领域的性能。数据集包含2,043,587条指令,涵盖了分子属性预测、蛋白质功能预测和生物分子文本理解等多个任务。创建过程中,数据从多个授权来源收集,并通过转换为适合特定任务的指令格式进行处理。该数据集的应用领域包括加速药物开发、揭示新的生物分子研究领域,并提升大型模型对生物学的理解能力。
Mol-Instructions is a large-scale biomolecular instruction dataset developed by the College of Computer Science and Technology, Zhejiang University. It aims to improve the performance of large language models (LLMs) in the biomolecular domain through three core components: molecule-oriented instructions, protein-oriented instructions, and biomolecular text instructions. The dataset contains 2,043,587 instruction entries, covering multiple tasks such as molecular property prediction, protein function prediction, and biomolecular text understanding. During its creation, data was collected from multiple authorized sources and processed into task-adapted instruction formats. The application scenarios of this dataset include accelerating drug development, uncovering new research directions in biomolecular fields, and enhancing the ability of large models to understand biological knowledge.




