silkome-full-idv-grouped
收藏资源简介:
Silkome Full Idv Grouped 是一个用于从丝蛋白组(Silkome)序列预测蜘蛛丝牵引丝纤维机械性能的分组数据集。该数据集由原始silkome-full数据集重构而来,核心创新在于将共享相同测量纤维/属性标识符(idv)的所有可用丝蛋白序列行聚合为单个样本,形成序列集合 -> 机械性能的映射关系,旨在服务于基于ESMC等嵌入模型和后续集合级聚合模型的预测流程。数据集包含270个唯一的idv分组样本,总计代表3563条蛋白质序列行。每个分组样本包含一个idv对应的所有丝蛋白氨基酸序列列表(sequences)、每条序列的类别标签(如MaSp1, MaSp2等,共18个类别)、序列长度以及基于类别的组成特征。目标变量是纤维级的四个关键机械性能:韧性(toughness)、杨氏模量(E)、强度(strength)和应变(strain),同时提供了这些性能的标准差及归一化版本。数据还附带了丰富的元数据,包括物种分类学信息(科、属、种)、性别、NCBI标识符等。数据集已预先划分为训练集(230个样本)和测试集(40个样本),采用基于property_tuple_key的确定性分组分割方法,确保训练集和测试集在idv和四元组属性上均无重叠,以进行更严谨、防泄漏的模型评估。该数据集适用于纤维级机械性能预测、序列嵌入与集合聚合模型基准测试、丝蛋白类别组成分析以及序列-性能关系研究等任务。需要注意的是,目标性能是纤维水平的测量值,并非单个蛋白质的直接功能标签,纤维力学还受多种非序列因素影响。
Silkome Full Idv Grouped is a grouped dataset for predicting the mechanical properties of spider dragline silk fibers from Silkome sequences. This dataset is reconstructed from the original silkome-full dataset. The core innovation lies in aggregating all available silk protein sequence rows that share the same measurement fiber/property identifier (idv) into a single sample, establishing a mapping relationship from sequence collections to mechanical properties, and it is designed to support prediction workflows based on embedding models such as ESMC and subsequent collection-level aggregation models. The dataset contains 270 unique idv-grouped samples, totaling 3563 protein sequence rows. Each grouped sample includes the list of all silk protein amino acid sequences (sequences) corresponding to one idv, the category label for each sequence (e.g., MaSp1, MaSp2, 18 categories in total), sequence lengths, and category-based compositional features. The target variables are four key fiber-level mechanical properties: toughness, Young's modulus (E), strength, and strain, with standard deviations and normalized versions of these properties also provided. The dataset also comes with rich metadata, including species taxonomic information (family, genus, species), gender, NCBI identifiers, and more. The dataset has been pre-divided into a training set (230 samples) and a test set (40 samples) using a deterministic grouping split method based on property_tuple_key, ensuring no overlap between the training and test sets in terms of idv and the four-tuple properties, enabling more rigorous and leakage-proof model evaluation. This dataset is applicable to tasks such as fiber-level mechanical property prediction, sequence embedding and collection aggregation model benchmarking, silk protein category compositional analysis, and sequence-property relationship research. It should be noted that the target properties are fiber-level measurements, not direct functional labels for individual proteins, and fiber mechanics are also influenced by a variety of non-sequence factors.




