遇见数据集

Dataset of 12,500 In Silico Mutant Enzyme Sequences from 13 Diverse Parents Generated by Monte Carlo Algorithm

收藏
Zenodo2025-06-19 更新2026-05-26 收录
官方服务:

资源简介:

Description: This dataset contains over 12,500 amino acid sequences representing in silico-generated mutant enzymes derived from 13 structurally and functionally diverse parent proteins. The parent enzymes range from 100 to 800 residues in length and exhibit minimal pairwise sequence similarity, ensuring broad coverage across sequence space. All variants were generated using a novel Monte Carlo-based mutagenesis framework, which applies stochastic local edits to amino acid sequences while preserving core structural features. Each mutant sequence was selected to maintain a full-protein backbone RMSD of less than 2 Å from its parent structure, ensuring structural plausibility and retaining functional fold integrity. Active sites for each enzyme were defined and preserved during sampling; RMSD values for active site residues are reported for all mutants. Mutation sampling was guided by tunable Monte Carlo parameters, including `mc_temp` and `mc_steps`, which control the exploration depth and energy threshold of the sequence edits. The algorithm supports substitutions, insertions, and deletions but prioritizes conservative mutations near active regions. All sequences are provided in JSON and FASTA formats, along with per-sequence metadata including: Parent PDB ID and class-level annotation (e.g., Oxidoreductase, Hydrolase, etc.) Mutation site positions and counts Global and active-site RMSD Sequence identity to parent Monte Carlo sampling parameters Example Entry Format: { "pdb_id": "1THX", "enzyme_class": "Oxidoreductase", "active_site_residues": ["W36", "C37", "G38", "P39", "C40"], "mutants": [ { "sequence": "RETPMSKGVITIKDAEF...", "sequence_identity": 0.91, "rmsd_active_site": 0.07, "rmsd_full_protein": 1.93, "plddt_active_residues": 0.865, "mc_params": { "mc_temp": 0.001, "mc_steps": 700, "mutation_energy": 0.01061 } } ] } Metadata: Number of mutants: 12,519 Number of parent enzymes: 13 Sequence length range: 100–800 residues Max full-protein RMSD: < 2.0 Å Format: JSON (with sequence metadata), FASTA License: CC BY 4.0 Contact: adam@aminoanalytica.com

描述: 本数据集包含超过12500条氨基酸序列,均为源自13种结构与功能各异的亲本蛋白的计算机模拟生成的突变酶。亲本酶的长度介于100至800个残基之间,且两两序列相似性极低,确保了序列空间的广泛覆盖。所有突变体均通过一种新颖的基于蒙特卡洛的诱变框架生成,该框架对氨基酸序列进行随机局部编辑,同时保留核心结构特征。 所有突变序列均经过筛选,确保其与亲本结构的全蛋白主链均方根偏差(Root Mean Square Deviation, RMSD)小于2埃,以保证结构合理性并维持功能折叠完整性。每一种酶的活性位点均已预先定义,并在采样过程中得到保留;所有突变体均报告了其活性位点残基的均方根偏差值。 突变采样由可调谐的蒙特卡洛(Monte Carlo)参数调控,其中`mc_temp`与`mc_steps`分别控制序列编辑的探索深度与能量阈值。该算法支持替换、插入与缺失突变,但优先在活性区域附近引入保守突变。 所有序列均以JSON与FASTA格式提供,并附带每条序列的元数据,具体包括: 1. 亲本PDB标识符(Protein Data Bank ID, PDB ID)与酶类注释(例如氧化还原酶、水解酶等) 2. 突变位点位置与数量 3. 全局与活性位点均方根偏差 4. 与亲本的序列一致性 5. 蒙特卡洛采样参数 示例条目格式: { "pdb_id": "1THX", "enzyme_class": "Oxidoreductase", "active_site_residues": ["W36", "C37", "G38", "P39", "C40"], "mutants": [ { "sequence": "RETPMSKGVITIKDAEF...", "sequence_identity": 0.91, "rmsd_active_site": 0.07, "rmsd_full_protein": 1.93, "plddt_active_residues": 0.865, "mc_params": { "mc_temp": 0.001, "mc_steps": 700, "mutation_energy": 0.01061 } } ] } 元数据: - 突变体总数:12519 - 亲本酶数量:13 - 序列长度范围:100~800个残基 - 最大全蛋白均方根偏差:<2.0埃 - 数据格式:JSON(附带序列元数据)、FASTA - 授权协议:CC BY 4.0 - 联系方式:adam@aminoanalytica.com

提供机构:
AminoAnalytica
创建时间:
2025-06-19
二维码
社区交流群
二维码
科研交流群
商业服务