lamm-mit/silkome-masp-idv-grouped
收藏资源简介:
Silkome MaSp Idv Grouped 是一个按 idv 分组、仅包含 MaSp 的 Silkome 数据集,用于从与相同测量的纤维/性能标识符相关联的主要壶腹丝素蛋白序列集合中预测拖丝纤维级别的机械性能。该数据集衍生自 lamm-mit/silkome-masp,不是将每个蛋白质序列放在单独的行中,而是将具有相同 idv 的序列行分组为一个样本:{MaSp sequences for one idv} -> {toughness, E, strength, strain}。这种设计旨在用于 ESMC 嵌入模型,后接集合级聚合模型。数据集包含 233 个唯一 idv 组,分为训练集(198 组)和测试集(35 组),无训练/测试 idv 或四属性元组重叠。源数据在过滤/分组前有 1033 行,每个 idv 的序列数最小为 1,中位数为 4.0,平均为 4.43,最大为 13。最大 X-linker 连接序列长度为 5896 个残基/令牌。类别计数包括 MaSp(223 行)、MaSp1(349 行)、MaSp2(331 行)、MaSp2B(40 行)、MaSp3(43 行)和 MaSp3B(47 行)。推荐模型输入包括 sequences、sequence_categories、sequence_lengths 和 category_counts_json 等列。目标列包括 toughness、E、strength、strain 及其标准化版本。数据集适用于纤维级机械性能预测、ESMC 嵌入加集合聚合基准测试、序列类别组成和序列-性能关系分析,以及比随机行级序列分割更安全的泄漏评估。
Silkome MaSp Idv Grouped is an idv-grouped MaSp-only Silkome dataset for predicting dragline fiber-level mechanical properties from the set of major ampullate spidroin sequences associated with the same measured fiber/property identifier. This dataset is derived from lamm-mit/silkome-masp. Instead of placing one protein sequence in each row, this dataset groups sequence rows with the same idv into one sample: {MaSp sequences for one idv} -> {toughness, E, strength, strain}. This formulation is designed for ESMC embedding models followed by a set-level aggregation model. The dataset contains 233 unique idv groups, split into train (198 groups) and test (35 groups) with no train/test idv or 4-property tuple overlap. Source rows before filtering/grouping are 1033, with sequences per idv: min 1, median 4.0, mean 4.43, max 13. Maximum X-linker concatenated sequence length is 5896 residues/tokens. Category counts include MaSp (223 rows), MaSp1 (349 rows), MaSp2 (331 rows), MaSp2B (40 rows), MaSp3 (43 rows), and MaSp3B (47 rows). Recommended model inputs include sequences, sequence_categories, sequence_lengths, and category_counts_json columns. Target columns include toughness, E, strength, strain, and their normalized versions. The dataset is intended for fiber-level mechanical property prediction, ESMC embedding plus set aggregation benchmarks, analysis of sequence category composition and sequence-property relationships, and leakage-safer train/test evaluation than random row-level sequence splits.




