遇见数据集

Foundation model representation-geometry analysis on a genomic testbed: 10-condition factorial dataset, 4-metric robustness analysis, and trained checkpoints

收藏
Zenodo2026-05-11 更新2026-05-26 收录
官方服务:

资源简介:

<p>Reproducibility package for the manuscript <em>"When do foundation model representations transcend trivial baselines? Training-objective design as the decisive factor"</em> (Tanigawa &amp; Iwaki, 2026, submitted).</p> <p><strong>Question.</strong> Foundation models are assumed to learn representations that transcend trivial statistical baselines of their input domain &mdash; <em>k</em>-mer composition in genomics, <em>n</em>-gram statistics in NLP, colour histograms in vision &mdash; but whether they actually do is rarely tested rigorously. Genomics offers a rare configuration in which a trivial baseline (<em>k</em>-mer composition), a ground-truth distance (evolutionary divergence times curated by literature consensus), and a residual statistic (partial Mantel) are all precisely defined. We exploit this configuration as a clean testbed for foundation-model representation evaluation, using splice sites as a controlled substrate.</p> <p><strong>Contents:</strong></p><ul><li>52,280 splice-site sequences from 140 species (39 Mammalia, 80 Insecta, 21 Nematoda) across 5 functional gene categories</li><li>10-condition factorial training-objective dataset on HyenaDNA-medium-160k (6.55 M parameters): factorial axes are loss (next-token / contrastive / classification / regression / adversarial), label resolution (phylum / species / category), label correctness, and initialisation</li><li>30 multi-seed fine-tuned model checkpoints (6 conditions &times; 5 seeds)</li><li>34 model embeddings (HDF5) on the held-out aging + DNA-repair evaluation set (45,080 sequences)</li><li>Residual Mantel test results (9,999 permutations, BH-FDR per phylum) for all conditions</li><li>Five-model GFM benchmark embeddings (Evo2 7B, NT-v2 500M, DNABERT-2 117M, DNABERT-S 117M, HyenaDNA-large 243M)</li><li>Four-metric representation-alignment robustness analysis: residual Mantel, representational similarity analysis (RSA Spearman), k-nearest-neighbour preservation @ k=5, k-NN phylogenetic-retrieval mean average precision @ k=5</li><li>22 Python scripts forming the end-to-end reproducible pipeline (~5 hours wall-clock on a single RTX 3090)</li><li>5 supplementary tables: gene lists, per-species N content, full 30-seed Mantel results, 5-model benchmark, per-seed 4-metric robustness</li></ul> <p><strong>Headline results:</strong></p><ul><li>Across five GFMs (6.55 M&ndash;7 B parameters), larger-scale zero-shot pretrained models do not reliably recover residual evolutionary geometry under this evaluation; only DNABERT-S (the contrastively-pretrained 117 M model) captures any beyond-composition signal.</li><li>Random-label contrastive perturbation alone recovers ~80% of the lift achievable with real species labels, indicating that bias-erasure of pretraining priors dominates label-derived content.</li><li>The species-fine pretrained "default" recipe is the worst Tier-1 design; four single-substitution alternatives (phylum-coarse contrastive, supervised species classification, distance regression, random-init species contrastive) all outperform it (paired &Delta; &ge; +0.045, p &le; 0.013, n = 5 seeds).</li><li>Within Tier 1, objectives optimise different geometry scales: phylum-coarse contrastive supervision preferentially improves global interspecies alignment (paired &Delta; = +0.066, p = 0.0001) while species-fine supervision better preserves local neighbourhoods &mdash; a global-vs-local geometry trade-off, robust across four representation-alignment assays.</li></ul> <p><strong>Reproducibility:</strong> the entire 30-run multi-seed verification can be reproduced on a single NVIDIA RTX 3090 in approximately 5 hours wall-clock. All scripts use deposit-relative paths and are released under the deposit's license; data and embeddings are CC BY 4.0.</p> <p><strong>Related manuscript:</strong> Tanigawa &amp; Iwaki (2026), submitted to Nature Machine Intelligence.</p>

提供机构:
Zenodo
创建时间:
2026-05-11
二维码
社区交流群
二维码
科研交流群
商业服务