Preprocessed Datasets for Interactive Exploration of DPCfam and DPCstruct Protein Domain Classifications
收藏资源简介:
This deposit provides all preprocessed input files required to deploy DPCexplorer (https://doi.org/10.5281/zenodo.20575268), a Django-based interactive portal for the exploration of protein domain classifications produced by the DPCfam and DPCstruct pipelines. The application will soon be accessible at https://dpcexplorer.areasciencepark.it/ and is fully reproducible from https://github.com/RitAreaSciencePark/dpc_fam_and_struct_webapp. Both DPCfam and DPCstruct apply the Density Peak Clustering (DPC) algorithm to automatically group millions of protein domains into evolutionary families called metaclusters: 81,384 sequence-based (DPCfam) and 28,246 structure-based (DPCstruct), without any manual curation. Despite being openly available on Zenodo, exploring these datasets directly requires parsing gigabyte-scale archives, including an 81 GB XML file for DPCfam, which presents a real technical barrier for most researchers. DPCexplorer removes that barrier by turning these static repositories into a searchable web application with paginated metadata tables, a 3D molecular viewer (PDBe-Molstar), and one-click downloads of per-metacluster biological files. The files in this deposit are cleaned, database-ready derivatives of the two original datasets: DPCfam (doi:10.5281/zenodo.6900559) and DPCstruct (doi:10.5281/zenodo.13334296). They were produced by a dedicated Python preprocessing pipeline run on the ORFEO HPC cluster; all notebooks and scripts are openly available in the linked GitHub repository. This deposit consists of four archives: dpcexplorer_csv.tar.gz (3.1 GB): PostgreSQL-ready CSV files for the 10 core database tables, covering DPC shared entities, DPCfam metacluster properties and seed sequences, AlphaFold representatives, DPCstruct metacluster properties and representative seed sequences, and CATH/SCOP structural annotations. dpcfam_mcid_seeds.tar.gz (1.7 GB): One FASTA file per DPCfam metacluster (81,384 files), covering both the Standard subset (size ≥ 50 seeds) and DPCfamB (25 ≤ size < 50 seeds). dpcstruct_mcid_pdbs.tar.gz (1.2 GB): One ZIP subfolder per DPCstruct metacluster (28,246 total), each containing the representative PDB structures loaded by the integrated 3D viewer. dpcstruct_mcid_seeds.tar.gz (6.9 MB): One FASTA file per DPCstruct metacluster (28,246 files), holding the sequences of the representative domains. This work was carried out during a Research Internship at the Laboratory of Data Engineering (LADE), Area Science Park, Trieste, Italy, as part of the MDMC Master's programme at SISSA, and was funded by the European Union - NextGenerationEU (NFFA-DI, cod. IR0000015 and EFC, cod. SSU2024-00002).



