Ising gene-content denoiser: code, final models, and training data for ancestral genome reconstruction
收藏资源简介:
Models, training data and code for Improved ancestral genome reconstruction using a learned gene-content grammar. The denoiser is an Ising model over the presence/absence of 4,789 COG gene families, fitted to 113,104 bacterial and archaeal genomes (one per species representative in GTDB r220). Reconstruction of a gene repertoire is treated as inference under that model: an iterative, gated mean-field relaxation with a reaction-field correction, an adaptive temperature, conditioning on module completeness over 419 functional modules, and a third-order attention head. Validation is by whole-phylum hold-out, so every test genome comes from a phylum unseen in training. This record contains the source code, the production model checkpoints for the LBCA and LACA reconstructions and the generalist denoiser (ten cross-validation splits each), the training and validation data, the leave-clade-out E. coli fine-tunes used for the divergence-time ladder. See README.md for the file-by-file guide and for which model produced which result. Source code is also at github.com/ssolo/gene-content-grammar. Licensed CC BY-NC 4.0. Training data derived from GTDB r220 (CC BY-SA 4.0) and the NCBI COG database (public domain); KEGG-derived content is not redistributed.



