AikiSol: per-protocol calibrated protein solubility prediction — Data and Model Weights
收藏资源简介:
Companion deposit for the manuscript 'Protein solubility prediction has been over-engineered and under-curated: a cluster-disjoint benchmark' (Mysore, 2026). Contains: (1) the 31,449-row license-clean subset of pure_solubility v1.1 used to train the released 3-seed AikiSol curve-head ensemble (parquet with v9_split column); (2) the 3-seed ensemble checkpoints (ESM-2 650M backbone + 5-output curve head); (3) per-cohort prediction CSVs and result JSONs reproducing every quantitative claim in the paper that depends on the redistributable training pool; (4) external benchmark cohort scoring artefacts (ProgSol-yeast, ProtSolM-eSol-2.2K, Foldit-117, PLM_Sol-EOI216) for AikiSol, NetSolP, and PLM_Sol; (5) one-MD5-per-row training-pool sequence list for cluster-overlap audits. Sources whose licenses prohibit redistribution (CC BY-NC-ND 4.0) or require explicit publisher permission (Springer Nature default) are excluded; see DATA_SOURCES.md for the inclusion/exclusion table. The manuscript also reports a cluster-disjoint 5-fold benchmark trained on a larger 84,809-row protocol-stratified pool that includes ~22K license-restricted rows; that pool is not redistributed here, and readers wishing to reproduce that exact benchmark must reconstruct it from the upstream sources cited in the SI. Code: https://github.com/aikium-public/aiki-sol



