遇见数据集

Aiki-Sol Dataset v0.2.0 — per-protocol calibrated protein solubility prediction (data, splits, checkpoints, results)

收藏
Zenodo2026-05-13 更新2026-05-26 收录
官方服务:

资源简介:

Companion deposit for the manuscript 'Protein solubility prediction is bottlenecked by training-pool curation, not architectural complexity' (Rajagopalan, Sharma Meda, Shastry & Mysore, 2026). Contains: (Apache 2.0 tier) the 147,574-row canonical Aiki-Sol Dataset training pool, the 5-fold cluster-disjoint splits at 25% identity, the released Aiki-Sol v2 deployment checkpoint trained on the full pool, five per-fold checkpoints, per-cohort prediction CSVs, and aggregated result JSONs. (CC-BY-NC-ND 4.0 research-tier extension) Aiki-Sol v2 research-tier checkpoint, trained on the Aiki-Sol Dataset extended to 229,349 rows with additional upstream sources under restrictive licences; the trained checkpoint is shipped under the most-restrictive upstream tier, the underlying research-tier training CSV is not redistributed and is instead documented via a per-source manifest for independent reconstruction. Architecture: single fine-tuned ESM-2 650M backbone with a 6-output head (5 per-stringency outputs for 3,000g/10min, 6,000g, 32,000g, eSol-21,600g/30min, 100,000g; plus a dedicated stringency-marginal output for binary-only-labelled inputs). Code: https://github.com/aikium-public/aiki-sol

提供机构:
Zenodo
创建时间:
2026-05-13
二维码
社区交流群
二维码
科研交流群
商业服务