introvoyz041/LinkLlama-cap50-train
收藏资源简介:
该文件是用于训练LinkLlama cap-50模型的监督微调语料库。每条数据是一个JSON对象,采用指令式布局。数据来源:母体结构来自ChEMBL数据库,分子被碎片化为片段-连接器-片段三元组,并计算了分子和连接器的属性及合理性启发式规则。应用了cap-50平衡方案,确保单个连接器SMILES在最终训练集中出现不超过50次,以减少频繁连接器的记忆。规模:平衡后约有160万条训练数据。文件格式为JSON Lines,字段遵循LinkLlama/Axolotl Alpaca-style约定。
This file is the supervised fine-tuning corpus used to train the LinkLlama cap-50 model. Each line is one JSON object in an instruction-style layout. Provenance: Parent structures were drawn from ChEMBL, molecules were fragmented into fragment–linker–fragment triplets, and molecular and linker properties and reasonability heuristics were computed. A cap-50 balancing scheme was applied so that no single linker SMILES appears more than 50 times in the final training set. Scale: On the order of ~1.6M training lines after balancing. Format: JSON Lines, with fields following the LinkLlama/Axolotl Alpaca-style convention.



