HyDB — an observation corpus and sequence registry for PET hydrolases
收藏资源简介:
HyDB is a provenance-tracked corpus of experimental observations on poly(ethylene terephthalate) (PET) hydrolases, together with the sequence registry the observations are anchored to. Nothing here is asserted: the package verifies itself (python3 VERIFY.py, standard library only, 18 checks). Contents (0.1.103), measured in the archive: 40,929 observations routed exhaustively and exclusively across five views — PET 28,094, MHET 1,317, SYS 5,056, CORE 6,289, AVAL 173; 12,147 observations at the direct-PET counter; 5,974 sequences of which 5,912 are redistributed in full and 62 withheld by explicit licence decision; 306 sources over 258 DOIs; 38 acquisition rounds, each locked by an invariant block. What changed in 0.1.103. The corpus is unchanged — no row admitted, none removed. Three things were repaired. Source titles were being lost on read. venue and year fell back to the registration-agency table; title did not. Fixing that one missing fallback recovered 68 source titles. Each title now carries a title_basis saying where it came from, so an empty field can be told apart from an unread one. The largest single source is now fully citable. The deposit behind 58 % of the sequence table carries a full title, fourteen authors with ORCIDs and a CC-BY-4.0 licence, and its DataCite type is JournalArticle — it is the article. Its source record now also carries what no agency holds: how the library was built (a random subset of 310 uncharacterised PETase genes plus site-saturation variant libraries on three scaffolds), the limit of quantification (35 µM MHET), and the variability its own authors report. Every observation from that source carries the assay platform that produced it. The study evaluates three platforms and recommends only one for model training; the label makes that distinction survivable downstream. Nothing is filtered — the authors' declared variability travels next to the label so the reader judges on their figure, not ours. Reproducibility. The two upstream invariant suites (94 blocks + 16) ship in proof/ and now run standalone: previously they only worked through the launcher and raised tracebacks when invoked directly, which made a folder named proof prove nothing. Concentration is measured, not corrected. Herfindahl indices and effective group counts per DOI and per sequence; three weight columns attached to every observation and none applied. The salient figure is the replication floor: the large majority of direct-PET sequences are seen by a single DOI. Derived from the PETHyDB curation deposit (10.5281/zenodo.21719676).



