遇见数据集

HyDB — an observation corpus and sequence registry for PET hydrolases

收藏
Zenodo2026-08-04 更新2026-08-13 收录
官方服务:

资源简介:

HyDB is a provenance-tracked corpus of experimental observations on poly(ethylene terephthalate) (PET) hydrolases, together with the sequence registry the observations are anchored to. Every value carries the source it was read from, the route it was extracted by, and the round it entered in. Nothing in this package is asserted: the package verifies itself. Contents of this version (0.1.102), as measured in the archive: 40,929 observations, routed exhaustively and exclusively across five views — PET (28,094), MHET (1,317), SYS (5,056), CORE (6,289), AVAL (173) 12,147 observations at the direct-PET counter: measurements on purified enzyme against PET polymer, in absolute units 5,974 sequences, of which 5,912 are redistributed in full and 62 are withheld by an explicit licence decision, their residues replaced by a sentinel and the record kept 306 sources over 258 distinct DOIs, and 37 acquisition rounds, each locked by an invariant block The direct-PET counter is derived, not declared. It is the conjunction of six orthogonal axes — target, level, detection, value_type, role, state — minus named demotions (duplicates, explicit do-not-count flags), and scoped to the part of the corpus normalised under the PETHyDB convention. The conjunction alone is necessary but not sufficient: applied naively it yields 13,934 rows instead of 12,147. The axes_convention column makes the scope explicit, and every demotion carries a readable motive. Concentration is measured and published, not corrected. Herfindahl indices and effective group counts are given per DOI and per sequence: 11.2 effective sources for the corpus by DOI against 76.1 by sequence; 6.9 and 44.3 respectively on direct-PET. Three weight columns (poids_inverse_doi, poids_inverse_sequence, sequence_multi_source) are attached to every observation and none is applied — filtering, stratifying and weighting are three different gestures, and the choice belongs to the consumer. The salient figure is not the dominant source but the replication floor: the large majority of direct-PET sequences are seen by a single DOI. Self-verification. Run python3 VERIFY.py at the root of the archive; it needs the standard library only. It checks properties rather than totals — a total copies from one document to another without anything being true. Fifteen checks cover manifest integrity, routing totality and exclusivity, the three counter properties, the AVAL boundary, redaction of non-redistributable residues, the residue alphabet, exactness of the concentration weights, and lineage completeness. proof/run_proofs.py replays the upstream invariant suites from the shipped snapshots. Licence and redaction. The package redistributes no article text. Sequences carrying an upstream licence that forbids redistribution are withheld by decision, not by omission: see LICENCES.md and DETTES.md, which also records the known debts of this version, including 34 chimeric records removed and 6 identifiers of declared uncertain provenance. Derived from the PETHyDB curation deposit (10.5281/zenodo.21719676), which holds the upstream tables and per-round reports.

提供机构:
Zenodo
创建时间:
2026-08-04
二维码
社区交流群
二维码
科研交流群
商业服务