遇见数据集

Ising gene-content denoiser: code, final models, and training data for ancestral genome reconstruction

收藏
Zenodo2026-09-03 更新2026-10-01 收录
官方服务:

资源简介:

Models, training data and code for Improved ancestral genome reconstruction using a learned gene-content grammar. The denoiser is an Ising model over the presence/absence of 4,789 COG gene families, fitted to 113,104 bacterial and archaeal genomes (one per species representative in GTDB r220). Reconstruction of a gene repertoire is treated as inference under that model: an iterative, gated mean-field relaxation with a reaction-field correction, an adaptive temperature, conditioning on module completeness over 419 functional modules, and a third-order attention head. Validation is by whole-phylum hold-out, so every test genome comes from a phylum unseen in training. This record contains the source code, the production model checkpoints for the LBCA and LACA reconstructions and the generalist denoiser (ten cross-validation splits each), the training and validation data, the leave-clade-out E. coli fine-tunes used for the divergence-time ladder. See README.md for the file-by-file guide and for which model produced which result. Source code is also at github.com/ssolo/gene-content-grammar. Licensed CC BY-NC 4.0. Training data derived from GTDB r220 (CC BY-SA 4.0) and the NCBI COG database (public domain); KEGG-derived content is not redistributed.

提供机构:
Zenodo
创建时间:
2026-09-03
二维码
社区交流群
二维码
科研交流群
商业服务