AikiSol v0.2.0 — per-protocol calibrated protein solubility prediction (data, splits, checkpoints, results)
收藏资源简介:
Companion deposit for the manuscript 'Protein solubility prediction is bottlenecked by training-pool curation, not architectural complexity' (Mysore et al., 2026). v0.2.0 supersedes v0.1.x. Contains: (Apache 2.0 tier) the 147,574-row canonical-clean training pool, the 5-fold cluster-disjoint splits at 25% identity, the released AikiSol-v2 deployment checkpoint trained on the full pool, five per-fold checkpoints, per-cohort prediction CSVs, and aggregated result JSONs. (CC-BY-NC-ND 4.0 research tier) AikiSol-v2-research, a checkpoint trained on the 229,349-row n3v2 pool that extends the canonical pool with additional upstream sources under restrictive licences; the trained checkpoint is shipped under the most-restrictive upstream tier, the underlying n3v2 training CSV is not redistributed and is instead documented via a per-source manifest for independent reconstruction. Architecture: single fine-tuned ESM-2 650M backbone with a 6-output head (5 per-stringency outputs for 3,000g/10min, 6,000g, 32,000g, eSol-21,600g/30min, 100,000g; plus a dedicated stringency-marginal output for binary-only-labelled inputs). Code: https://github.com/aikium-public/aiki-sol



