遇见数据集

PANTS: a triaged catalogue of candidate PET-degrading enzymes for therapeutic use

收藏
Zenodo2026-08-08 更新2026-08-13 收录
官方服务:

资源简介:

PANTS mines metagenomic sequence space for polyester hydrolases and triages them for therapeutic use: degrading PET at 37 °C in serum rather than in an industrial reactor above PET's glass transition. Version 0.3.0 adds the first measured negative class this field has had, and the three findings it made possible. 14,804,920 predicted proteins scanned across five environments, 439 candidates retained, 416 with a predicted structure. The reference set holds 1348 characterised enzymes and 2,759 activity measurements (2,757 carrying a DOI), with 1,395 structures deposited alongside. New in 0.3.0: 150 measured negatives. Every earlier evaluation rested on negatives meaning "not reported active in a database", which cannot distinguish an enzyme assayed and found dead from one nobody assayed. Two 2025 high-throughput screens — ACS Catalysis (10.1021/acscatal.5c03460) and Science (10.1126/science.adp5637) — assayed panels under single protocols and reported what did not work. Their openly deposited source data supplies enzymes expressed, assayed, and found to release no product. Three negative results follow, and they are the substance of this release. 1. Active-site geometry does not predict PET activity. On inferred labels cleft depth separated the classes at AUC 0.808, p 1.7×10−7. On measured labels, 237 active against 139 inactive, all ESMFold on both sides: AUC 0.398. The trajectory as negatives accumulated — 0.507, 0.498, 0.459, 0.398 — moves toward chance as the sample grows, which is a confound diluting rather than an underpowered test. 2. The determinants are lineage-specific, which explains why nothing transfers. Fitting the same model separately inside each large lineage and comparing directions: two fits of the same lineage agree at +0.73, but different lineages agree at -0.32 to +0.34, and in two of three pairs point in opposite directions. The same pattern appears in two feature sets sharing nothing, one from 3D coordinates and one from a protein language model. A model trained across lineages learns a direction that reverses on a lineage it has not seen. 3. Retrieval is the honest deliverable. A learned head is at chance across lineages and cannot be shown to beat nearest-neighbour retrieval within one (paired lead +0.159 [−0.263, +0.390]). Eighteen times the language-model parameters changed nothing. Candidates are therefore ranked by identity to the nearest enzyme whose activity was measured, each carrying a competence band from the measured identity-decay curve: 16 in range, 39 marginal, 384 out of range. The 384 out-of-range candidates are precisely the novel enzymes this project set out to find, and nothing here can score them — stated on the site rather than hidden. Every number was recomputed from the deposited files by scripts/build_release.py, never copied from prose, and each analysis ships its own JSON artefact. Data CC BY 4.0; source code MIT.

提供机构:
Zenodo
创建时间:
2026-08-05
二维码
社区交流群
二维码
科研交流群
商业服务