遇见数据集

Substrate Site Discovery — repository snapshot 2026-09-20 (option-C redacted, revision 2)

收藏
Zenodo2026-09-24 更新2026-10-01 收录
官方服务:

资源简介:

What this is This record publishes a development workspace snapshot of an ongoing project on substrate-conditioned enzyme active-site discovery, together with the benchmark material and the governance/evidence record that the workspace accumulated between 2026-09-08 and 2026-09-20. This is engineering-prototype and evidence material. It is not a set of scientific conclusions. The workspace's own standing rule is that its outputs are prototypes and drafts that do not replace any formal conclusion, and that no enzyme activity may be inferred from its staging or governance files and none of it may be used for training. That rule is carried through verbatim into this record. Contents 1,610 files; 17,212,330 bytes compressed, 79,394,320 bytes uncompressed. Of these, 21 files carry annotated redactions and a further 17 carry hash-cross-reference annotations (see below). - Code (`discovery/`, `tools/`, `bench/`, `schema/`): candidate-region discovery, a catalytic-rule engine, a substrate-to-candidate ranker, a tiered benchmark registry and report tooling, and a `site.json` schema v2 validator. Released under the Apache License 2.0 — see `LICENSE`. - Data and text (`corpus/`, `substrate_spec/`, `truth_table/`, `positive_standard/`, `inventory/`, `docs/`, `bench/` CSVs, root-level tables): released under CC-BY-4.0 — see `LICENSE-DATA`. - Governance and evidence record (`governance/`, 159 files): the project's own audit machinery and its outputs — an activity-truth ledger schema and validator, checker censuses, numbering and residue-map audits, register-geometry and sealed-negative audits, quote-provenance verification, a derived-artifact freshness check, three-state wording checks, repository-hygiene audits, and offline-anchor registries. Scope and honest limitations - No admissible activity ground truth. The activity-truth ledger's `records` array is empty and `n_activity_records` is 0. A candidate set of 13 records migrated from an existing kinetic table is held as _staging_: all 13 are classified `未判定` (undetermined) because they lack verifiable assay method and independent replicate counts; 7 also lack uncertainty. The workspace's stated position is that structure/contact candidate generation may be developed at this stage, but it may not be claimed that enzyme activity can be inferred from substrate. - Truth table pending expert sign-off. The reaction/substrate/cleavage-site truth table is a draft; a subset of its rows is approved at _EC-reaction level only_, with the remaining flags left open. - Benchmark sizes are small and partly ungraded. Tier assignments were only partly resolved; where a tier could not be assigned the workspace reports it as empty rather than filling it in. Reported recall figures come from very small panels and should be read with that in mind. - Reported negatives are reported as negatives. Where a component degraded performance (for example on two of the panel targets) the workspace records that plainly rather than omitting it. Snapshot provenance (please read) The archive is a reconstruction, not a copy of a single directory. At packaging time the working copy at `priority_work/substrate_site_discovery` had lost its git metadata directory (present but empty) and retained only 258 files; its project subdirectories were empty while its governance files continued to be updated through 2026-09-20. The complete version-controlled tree (1,580 files, umbrella repository commit `c5eabc5`, 2026-09-17) was combined with the newer governance files (26 added, 19 updated) from the working copy. Per-file provenance for every member is recorded in `PACKAGE_PROVENANCE.tsv`, distributed as a separate file in this record. Nine root-level files that existed only in the working copy were excluded: they are stale copies of paths that no longer exist in the 2026-09-17 tree. The archive carries `LICENSE` (Apache-2.0), `LICENSE-DATA` (CC-BY-4.0) and `NOTICE` at its root. These did not exist anywhere in the project before this packaging; they were added to the archive only and have not been written into any working tree. Annotated redaction of compute-infrastructure identifiers 21 documents carry an annotated redaction: the hostname, instance ports and container identifiers of the rental GPU instances used for structure prediction were replaced with `[COMPUTE-HOST-REDACTED]`, `[COMPUTE-PORT-REDACTED]` and `[COMPUTE-CONTAINER-REDACTED]`, and each affected file carries an appended erratum notice recording the SHA-256 of its unmodified original. Nothing else in those files was altered. Because this project's governance discipline requires every artifact's hash to be recorded in other documents, redacting those 21 files invalidated hash and byte-size cross-references elsewhere. A further 17 documents therefore carry an appended cross-reference annotation stating that the hash and size values they record refer to the pre-packaging originals (for `.json` this is an added top-level key `_c_prime_annotation`; for the one machine-readable `.tsv` manifest it is a single comment line appended at end-of-file, leaving every existing row untouched). The transitive closure of this cascade was computed over four rounds and verified to terminate — there is no document that became stale without also being annotated. The complete record — all 38 paths, the categories replaced, and every original SHA-256 — is in `REDACTION_NOTICE.md` at the archive root, together with the full list of the 29 now-superseded cross-references across 19 files. The source working trees were not modified. No password, key, or token is present anywhere in the archive, before or after redaction; the project's own discipline keeps credential values out of all documents and logs, and an independent three-pass scan confirmed this. Two files under `corpus/` contain a digit string identical to one of the instance ports, but in context it is a PDB atom serial number — a coincidental match with the scanner's port pattern. Those two files were deliberately not altered, and the notice intentionally does not restate the digits. Licensing and rights Code is released under the Apache License 2.0; data and text under Creative Commons Attribution 4.0 International (CC-BY-4.0). Copyright (c) 2026 Yao, Jian; Shi, Chenyang; Li, Zhiyi. Upstream third-party material (RCSB PDB, UniProt, PDBe, published literature, Python packages) remains under its own terms and is not re-licensed here. Upstream model parameters are not distributed with this snapshot. Related records This is an independent concept record. It is related to, but is not a version of, the sibling project's record `10.5281/zenodo.22545060`. The two are parallel projects in the same programme; do not treat one as a version of the other. Integrity Archive SHA-256: `c0db48a03ff7db668a7fe86d0c4dd4ed221d0440d7e36deefff4fcc9ac800a52`; MD5: `95076816224393e7e7fc7d4a62563a24`. Per-file hashes are in `MANIFEST_sha256.tsv`, distributed as a separate file in this record rather than inside the archive, so that it can list every member without a self-reference. Extract the archive and recompute: 1,610 / 1,610 match. Known stale in-package statement This record's DOI has since been assigned: 10.5281/zenodo.22939366 (concept DOI 10.5281/zenodo.22939365). One statement inside the archive predates that assignment: LICENSE-DATA, section 3, still reads that the record's DOI is not yet allocated and should be back-filled on publication. That sentence is part of the packaged bytes and is therefore left unmodified here, following this project's standing rule that dated snapshots are registered rather than rewritten. It is superseded by this record.

提供机构:
Zenodo
创建时间:
2026-09-24
二维码
社区交流群
二维码
科研交流群
商业服务