protease_plm_structure_corrective
收藏资源简介:
This archive accompanies the manuscript “Three-Dimensional Structure as a Conditional Corrective in Protease Annotation by Protein Language Models”. It contains the data subsets, model outputs, and scripts used to build and analyze a sequence–structure benchmark for protease-like proteins. Raw sequences are sampled from MEROPS v12.5 and UniProtKB/Swiss-Prot, and structures are taken from the AlphaFold Protein Structure Database. Folder overview raw_data/ — Sampled raw FASTA files for protease and non-enzyme sequences. processed_data/ — Processed benchmark subsets, including train/validation/test sequence splits (FASTA), AlphaFold PDB structures, residue–contact graphs (.pt), and metadata tables. results/ — Outputs from all models (ESM sequence-only, GNN structure-only, fusion and ablation variants), including predictions and performance metrics. figure_source_tables/ — Source CSV tables for all figures in the manuscript (per-protein test-set summaries and plotting data). scripts/ — Main Python scripts for dataset construction, model training, evaluation, ablation analysis, and feature extraction.



