trillionlabs/TheBioCollection
收藏资源简介:
TheBioCollection是一个包含526亿令牌的预训练规模生物学语料库,旨在将异构生物资源转化为适合大语言模型训练的数据。它通过一个构建流程创建,该流程收集跨生物领域的资源,通过去重、实体标记和增强进行精炼,用工具计算的生物学属性进行丰富,并将其呈现为具有可编程验证答案的指令形式数据。该语料库涵盖广泛的生物领域,包括小分子、蛋白质、基因组序列、细胞和通路。
TheBioCollection is a 52.6B-token pretraining-scale corpus for biology that transforms heterogeneous biological resources into LLM training-friendly data. It is built through a construction pipeline that collects resources across biological domains, refines them through deduplication, entity tagging and augmentation, enriches them with tool-computed biological properties, and render them as instruction-form data with programmatically verifiable answers. The corpus spans broad biological domain across small molecules, proteins, genomic sequences, cells, and pathways.




