遇见数据集

HyDB — an observation corpus and sequence registry for PET hydrolases (v1.0)

收藏
Zenodo2026-08-04 更新2026-08-13 收录
官方服务:

资源简介:

HyDB gathers into a single table of record the published activity measurements for enzymes that degrade poly(ethylene terephthalate) and its intermediates, together with a registry of the protein sequences those measurements attach to. PETHyDB and MHETHyDB are two readings of it, obtained by selection rather than by parallel curation. 38,269 observations across 228 DOIs, 2,850 sequence identities addressed by content, 2,172 domains annotated by computation. Five rules govern the whole, each enforced by code and watched by an invariant: the upstream table is read-only and its digest is checked before and after every build; routing is a total function, so no row is dropped for want of a destination; counters are derived and never asserted, and any decrease requires a declared rebase with a written reason; the identity of a sequence is the digest of its residues, not its name; and an absence is classified — not applicable, debt, declared refusal, or out of reach — instead of being conflated with the others. What the resource does not know is published on the same footing as what it knows: 7,823 rows without pH, 4,555 without a link to a sequence, and 813 rows refused from two acquisition packages whose extraction produced values absent from the sources — each with its class, its reason, and the context excerpt that allows the refusal to be contested. The deposit is reproducible: run_all.sh rebuilds everything from the upstream tables and replays 34 invariants. No output is hand-edited. Note: the repository's documentation, code comments and refusal reasons are written in French.

提供机构:
Zenodo
创建时间:
2026-08-04
二维码
社区交流群
二维码
科研交流群
商业服务