遇见数据集

Constructing the ensemble of representative structures for a protein via neural-surrogate-guided MSA recombination

收藏
Zenodo2026-01-14 更新2026-05-26 收录
官方服务:

资源简介:

ProCEDiS – Data package This record provides the datasets and analysis artifacts used in the ProCEDiS study.Files are distributed as ZIP archives for convenient download and sharing. After downloading, please unzip and keep the folder structure unchanged. For source code and detailed instructions on usage, please refer to our github . scores.zip Contains evaluation outputs (TM-score–based) for all methods. After unzipping, each method has one folder, and inside each method folder each PDB-ID subfolder provides score_list.csv (TM-scores of every generated structure against each reference label) plus representative-set index files rep_*_indices*.npy that map directly to the row indices of score_list.csv (for ProCEDiS variants, the extra suffix indicates the episode). dataset_77proteins.zip Contains the 77-protein benchmark inputs, including per-target sequences (.fasta) and MSAs (.a3m), plus ground-truth conformational pairs from the PDB under pair/. For analyses that require matched lengths, homologous_modeling/ provides length-aligned reference models generated by homology modeling. procedis_sample.zip Provides sample ProCEDiS outputs for the 77-protein benchmark (standard and high-temperature variants). For each target, msa_pool/ contains all MSAs explored during the search and structure_pool/ contains all corresponding predicted structures—both are raw, non-deduplicated pools. parallel_md_npy.zip Contains the parallel MD simulation outputs for selected systems and seeds: each seed folder includes the minimized system structure (*_minimized.cif), equilibration and production logs (*_npt_eq.log, *_npt_prod.log), and the production trajectory saved as Cα-only coordinates (*_npt_prod_prot_ca.npy) to reduce file size. ncbi_tax_species.zip Contains a single taxonomy lineage table (ncbi_tax_species.csv) that maps each NCBI tax_id to major ranks (superkingdom→species), used for taxonomic annotation in our Taxo-seq clustering. It was generated from the NCBI taxonomy files using the ncbitax2lin workflow (GitHub repo: zyxue/ncbitax2lin).

提供机构:
Zenodo
创建时间:
2026-01-14
二维码
社区交流群
二维码
科研交流群
商业服务