Supplementary materials for "MemConverter: An Iterative Pipeline for Reprogramming Protein Localization in Membrane or Aqueous Solution"
收藏资源简介:
Description This dataset contains 85,051 isolated transmembrane domains derived from the tmAFDB and TED. It was specifically curated to fine-tune ProteinMPNN for membrane protein design (MemProtMPNN). We focused on isolated domains to ensure the model learns specific membrane topological constraints rather than soluble features. Data Construction Domain Extraction: Primarily based on domain annotations from The Encyclopedia of Domains (TED). For proteins without TED annotations, Merizo was used to identify domains. Filtering: Validated by TMbed. We excluded proteins with <20% transmembrane residues or domains shorter than 32 residues. Splitting: Clustered using MMseqs2 (30% sequence identity) to strictly separate Training, Validation, and Test sets. Statistics Set Count Percentage Train 72,293 85% Validation 7,655 9% Test 5,103 6% Files Description membrane_domain_dataset.zip: The main dataset archive. data_files.zip: Contains the processed dataset partitions (Train, Validation, Test). test_list: A curated list of representative cases from the test set (one representative sequence per cluster) used for performance evaluation.
数据集概述 本数据集包含源自tmAFDB与TED的85051个独立跨膜结构域(transmembrane domains),系为膜蛋白设计场景下的ProteinMPNN微调任务专门整理构建,该微调模型命名为MemProtMPNN。 我们选取独立跨膜结构域作为研究对象,旨在确保模型能够学习膜蛋白特有的拓扑约束特性,而非可溶性蛋白的相关特征。 数据构建流程 结构域提取:主要基于结构域百科(The Encyclopedia of Domains,简称TED)的结构域注释信息进行提取;对于未获取TED注释的蛋白质,则通过Merizo工具识别其结构域。 筛选流程:所有数据均通过TMbed工具进行验证,我们剔除了跨膜残基占比低于20%的蛋白质,以及长度不足32个氨基酸残基的结构域。 数据集划分:采用MMseqs2工具,以30%序列同一性阈值完成聚类,严格划分训练集、验证集与测试集。 数据集统计 | 数据集划分 | 样本数量 | 占比 | | :--------- | :------- | :--- | | 训练集 | 72293 | 85% | | 验证集 | 7655 | 9% | | 测试集 | 5103 | 6% | 文件说明 1. membrane_domain_dataset.zip:主数据集归档文件 2. data_files.zip:包含已完成划分的训练集、验证集、测试集子集数据 3. test_list:从测试集中精选出的代表性样本列表(每个聚类仅保留一条代表性序列),用于模型性能评估



