Aiki-XP: leakage-controlled multimodal prediction of protein expression at pan-bacterial scale — Data and Model Weights
收藏资源简介:
Training data (492,026 bacterial genes, 385 species), pre-extracted foundation-model embeddings, trained 5-fold-ensemble checkpoints for five deployment tiers (A, B, B+, C, D), gene-operon and species-cluster split files, external benchmark embeddings, and result JSONs for the Aiki-XP paper. Reviewer reproduction entry point: tier_predictions_lookup.parquet (244,002 rows × 10 columns) — held-out 5-fold CV predictions for all test-split genes with gene_id, species, is_mega, cv_fold, true_expression, and the five tier predictions (tier_a through tier_d). Group by cv_fold and average Spearman ρ across folds to reproduce the manuscript’s headline numbers (e.g. Tier D ρnon-mega = 0.590 ± 0.012 with is_mega=False). Code: https://github.com/aikium-public/aiki-xp. Live demo: https://aikium--aikixp-tier-a-landing-page.modal.run.



