遇见数据集

Data archive: compressible-design-laws

收藏
Zenodo2026-07-07 更新2026-08-02 收录
官方服务:

资源简介:

This folder is the data payload for the code release accompanying Xu, Tu and Xu, "Effective dimensionality governs when combinatorial design landscapes compress into interpretable laws". It contains the datasets fetched by scripts/fetch_data.py / make data in the code repository. Layout: data/cleaned/ holds 13 curated design-to-phenotype datasets under a uniform schema (manifest.csv, SCHEMA.md, and the individual CSV tables including beta-carotene, tryptophan, limonene, isoprenol, jervis-RBS, pca-paperA, vanlent-pCA, smanski-nif, poelwijk, gb1, avgfp, aav, and the vanlent-simulator enumerated 6^7 simulator ground truth). data/proteingym/ holds proteingym_ref.csv and the DMS_ProteinGym_substitutions/ folder containing the 69 combinatorial (multi-mutant) DMS assays used (columns used: mutant, DMS_score). Provenance: Curated cleaned datasets were reduced to a uniform schema from the originally published sources listed (with DOI/accession) in cleaned/manifest.csv. Each table ends in a single continuous phenotype column; replicate designs are aggregated to their mean (see cleaned/SCHEMA.md). For ProteinGym, only the 69 assays with includes_multiple_mutants = true in proteingym_ref.csv are included here (the combinatorial subset the population analysis uses). The full 217-assay substitution benchmark is available from the official ProteinGym Zenodo record 10.5281/zenodo.15293562; this folder is a faithful subset of it, not a re-derivation. The code subsamples each assay to at most 1,500 designs at runtime (seed-locked), so the raw assay files are kept unmodified to preserve exact reproducibility. Provenance for every dataset is in cleaned/manifest.csv.

提供机构:
Zenodo
创建时间:
2026-07-07
二维码
社区交流群
二维码
科研交流群
商业服务