PET-Prob Public Evidence Corpus v0.3.0
收藏资源简介:
The PET-Prob Public Evidence Corpus v0.3.0 is an auditable, public-data-only corpus of experimentally characterized PET-hydrolase sequences for sequence-based machine learning (built as data construction only — no model was trained, tuned or evaluated). Contents: 869 gold/within-study canonical sequences with an evidence-tier ledger (A1 gold-direct-binary 51; A2 gold-direct-quantitative 371 = Beckham + VenusMine + Seo; B1 within-study-relative 447 = PETML), 422 binary-eligible sequences (275 active / 147 inactive); a 57,720-sequence public unlabelled cutinase/polyesterase domain corpus; four model-ready task tables (binary competence, condition activity, within-study ranking, weak supervision); five leakage-controlled split candidates (no final split chosen); MMseqs2 clustering; reconciliations; and full provenance / licence / public-source registries. Redistributable subset. This archive contains only legally redistributable data (CC-BY / MIT / NCBI-public sequences) plus metadata, checksums and scripts for public but non-redistributable sources (Seo Science SI, Glacier). Restricted raw files are excluded; their DOIs and SHA-256 are recorded for reacquisition. Every packaged file is listed in MANIFEST.sha256. Sources: Erickson 2022 (10.1038/s41467-022-35237-x), Norton-Baker/Beckham 2025 (10.1021/acscatal.5c03460), PETML, VenusMine (10.1038/s41467-025-61599-z), Seo 2025 (10.1126/science.adp5637; observed public development, not blind), and public catalogues (PAZy, PANDA, PlasticEnz, PlasticDB, PEPIC) as weak-supervision tiers.



