HyDB — an observation corpus and sequence registry for PET hydrolases
收藏资源简介:
HyDB is a provenance-tracked corpus of experimental observations on poly(ethylene terephthalate) (PET) hydrolases, together with the sequence registry the observations are anchored to. Nothing here is asserted: the package verifies itself (python3 VERIFY.py, standard library only, 18 property checks), and the two upstream invariant suites (95 blocks + 16) ship in proof/ and run standalone. Contents (0.1.104), measured in the archive: 40,929 observations routed exhaustively and exclusively across five views — PET 28,094, MHET 1,317, SYS 5,056, CORE 6,289, AVAL 173; 12,147 observations at the direct-PET counter, 85.0 % anchored to a sequence and 81.7 % carrying both temperature and pH; 5,974 sequences of which 5,891 are redistributed in full and 62 withheld by explicit licence decision; 306 sources over 258 DOIs; 39 acquisition rounds, each locked by an invariant block. The direct-PET counter is derived, not declared. It is the conjunction of six orthogonal axes — target, level, detection, value_type, role, state — minus named demotions, scoped to the part of the corpus normalised under the PETHyDB convention. Applying the documented conjunction naively yields 13,934 rows instead of 12,147: that 15 % gap is the measure of what an ordinary database would over-count without any single value being wrong. What changed in 0.1.104 No row admitted, none removed. Three things were established. Citability rose from 54 % to 93 % (284 of 306 sources now carry a title). The blocker was not missing data but a refusal of principle: the sources table fabricated "neither title nor year" for acquisition-extension sources. The principle holds, its application conflated two gestures. Guessing a title — reconstructing it from a filename — remains forbidden. Resolving one means asking the registration agency what it recorded for that identifier: that is the authority record, the very thing that makes a DOI citable. Every title now carries a title_basis naming its agency, and couverture stays EXTENSION_NON_CURATEE — knowing who published an article says nothing about whether the values drawn from it were verified. Inter-team reproducibility is not measurable on this corpus, and that is the result. Requiring the same sequence (by SHA, not by name), the same quantity, the same unit and two distinct DOIs yields 17 comparable groups over 11 sequences — and zero at strictly identical conditions. The 955× median spread is therefore confounded with the range of temperatures and pH values and cannot be attributed to any team. A watch entry now waits for the first genuinely controlled pair. A checksum authenticates bytes, not meaning. Four sequence records received this round were declared OBSERVED_EXACT with exact SHA-256, held lengths and valid alphabets — and all four are concatenations of several constructs. Alphabet, length and checksum all pass on a concatenation; only a structural check sees it. One panel decomposes with proof into 4 constructs; two nucleotide records fall to something simpler than any decomposition — 3,350 and 1,919 bases, neither a multiple of three. Concentration is measured, not corrected. Herfindahl indices and effective group counts per DOI and per sequence (11.2 vs 76.1 effective groups on the corpus; 6.9 vs 44.3 on direct-PET), three weight columns attached to every observation, none applied. The salient figure remains the replication floor: 91 % of direct-PET sequences are seen by a single DOI. Known debts are listed in DETTES.md, foremost that the extraction vocabulary does not yet recognise kinetic parameters (Km, kcat, kcat/Km, Tm), leaving 80 cells of real kinetic tables unread in four sources already in the corpus. Derived from the PETHyDB curation deposit (10.5281/zenodo.21719676).



