遇见数据集

Structure-guided extremophile genome mining expands the PETase landscape and reveals PET-hydrolysing true lipase lineages

收藏
Zenodo2026-08-13 更新2026-08-20 收录
官方服务:

资源简介:

Poly(ethylene terephthalate) (PET) hydrolases are enzymes primarily within the polyesterase–cutinase branch of the α/β-hydrolase superfamily, whereas the contribution of true lipases to PET hydrolysis remains poorly explored. Here, we report the genome mining results of 18,082 extremophilic microorganisms, combined with structural modelling, enzyme–substrate simulations, and experimental validation, which enables the identification of PET-active true lipases. Two true lipases, LipBv and LipSh1, hydrolyse PET substrates, and homologous enzymes within their respective clusters also retain PET-hydrolysing activity, supporting the existence of specific lipase lineages associated with PET hydrolysis. Structural analyses suggest that the PETase-like activity and hydrolysis product profiles of these lipases are associated with differences in loop organization and active site accessibility. These lipase lineages clustered separately from 1,322 putative PETase-like hydrolases from the same extremophile genomes, while all groups remained distinct from previously characterized PETases. These findings expand the evolutionary diversity of PET-degrading enzymes in extremophiles. All files associated with this dataset can be downloaded directly from the Zenodo repository. Spreadsheet files can be opened with Microsoft Excel, LibreOffice Calc, or compatible software. Raw mass spectrometry data require MassLynx V4.1 (Waters) for processing. Network files can be imported into Cytoscape. File names and numbering correspond to the supplementary materials cited in the associated publication. Table S1. Reference and candidate true lipases identified during the genome mining workflow. This table summarizes the datasets used to identify candidate true lipases from extremophilic genomes. Reference lipase sequences were compiled and these sequences were used as queries for similarity searches against the Genome Taxonomy Database (GTDB) and other protein databases. a) Reference dataset of true lipases compiled to be used as reference queries for BLAST searches. b) List of genomes retrieved from the GTDB selected possible extremophiles to search lipases. c) Initial list of selected sequences from extremophiles with at least 60% of identity percentage and minimum alignment length of 100 residues with the reference lipases. d) Filtered list of 50 presumptive true lipases identified, after signal peptide detection and reducing sequence redundancy, further preselected candidates for structural and functional analysis. Metadata and sequence features of candidate enzymes identified in this study, including GTDB genome IDs, UniProt matches, sequence identity, signal peptide prediction, sequences used for gene synthesis, environmental origin, taxonomic classification, catalytic triad residues, and predicted structural features. e) Final selected candidates after structural and functional analysis. Sequence features of candidate enzymes identified in this study, including GTDB genome IDs, UniProt matches, sequence identity and aligment length. It shows original and synthetic protein sequences (signal peptides in blue, SignalP-5.0), pET-45(+) plasmid constructs synthesized by GenScript Biotech (EG Rijswijk, The Netherlands; ref. U063CUMTG0, U790YVJEG0 and U5150FMPG0), antibiotic resistance (ampicillin), inducer (IPTG), molecular weights (calculated with ExPASy; https://web.expasy.org/compute_pi/), catalytic triad residues, predicted structural features, environmental origin and extremophily and taxonomic classification. Due to the size and complexity of the dataset, Table S1 is provided as a separate Microsoft Excel (.xlsx) file. All worksheets are labelled according to the workflow stages described in the manuscript and can be accessed using standard spreadsheet software. Table S2. Reference PETases and candidate PETase-like sequences identified from extremophilic genomes. This table summarizes the datasets used to identify putative PETase-like enzymes from extremophilic microbial genomes. Experimentally characterized PETase sequences were compiled and used as queries for homology-based searches against the GTDB. a) Reference dataset of experimentally validated PETases used as queries for similarity searches. b) Set of 1,322 candidate PETase-like protein sequences identified from 18,082 extremophilic microbial genomes retrieved from the GTDB. Metadata and sequence features are provided for each candidate, including GTDB genome ID, UniProt matches, sequence identity, taxonomic classification, and environmental origin. c) List of sequences for SSN described by Seo et al. (2025)1. Owing to the large size of the dataset, Table S2 is provided as a separate Microsoft Excel (.xlsx) file. The file contains multiple worksheets corresponding to the different stages of the PETase mining workflow and can be viewed using standard spreadsheet software. Data S1. Raw MS data. The data in the raw folder can be analyzed with the software MassLynx V4.1 (Waters) to obtain the mass spectra of PET degradation oligomers. The mass spectrometry files can be analyzed with the software MassLynx V4.1. Data S2. Unprocessed HPLC data. Files are provided in .txt format corresponding to the data extracted with Varian ProStar software (Varian Inc., Palo Alto, California, USA). Data S3. Network. Net file that reconstructs the network. The network file is provided in .net format and can be directly imported into Cytoscape (version 3.0 or later recommended) for visualization and further network analysis. Data S4. Raw data for in vitro tests. a) Lipase activity. Datasets include the absorbance at 550 nm using NEFA kit and the corresponding specific activity values after incubation of different triglycerydes, oils and esters with all of the purified enzymes. Data are provided in triplicates. Related to Table 3 and Table 4. b) PETase activity. Datasets include concentrations of degradation products and the corresponding specific activity values obtained after incubation of nPET and pPET with purified enzymes at 40 and 60 ºC. Quantification was performed by HPLC using calibration curves (original chromatograms in Data S2). Data are provided in triplicates, with both raw and processed values (including applied dilutions or concentration factors). Related to Table 4 and Table 5. c) Physicochemical characterization of LipSh1 and LipBv. Datasets include the absorbance at 550 nm using glyceryl trioctanoate and NEFA kit (temperature) or at 405 nm using p-nitrophenyl butyrate (pH) as substrates and the corresponding relative activity values after incubation. Data are provided in triplicates. Related to Figure 3. Raw and processed datasets are provided in spreadsheet format, including replicate measurements and calculated values used for data analysis. These files enable full reproduction of the enzymatic activity calculations and figures reported in the manuscript.

提供机构:
Zenodo
创建时间:
2026-08-13
二维码
社区交流群
二维码
科研交流群
商业服务