Clifti-GPT benchmark AnnData: reference and query splits for federated single-cell experiments
收藏资源简介:
Clifti-GPT benchmark AnnData Prepared reference and query AnnData (.h5ad) files used to reproduce Clifti-GPT experiments across six public single-cell cohorts. Each archive contains the two files required by the experiment registry in the Clifti-GPT repository. Gene expression is filtered to the scGPT vocabulary; highly variable genes (HVG) are not used for subsetting. Related deposit (init weights, separate record): 10.5281/zenodo.20489646 This deposit provides the minimal benchmark bundle for reproducing Clifti-GPT fine-tuning, annotation, and embedding experiments. Files are split into reference (labeled training) and query (held-out test) AnnData objects per cohort, matching the layout expected under data/scgpt/benchmark/ in the Clifti-GPT codebase. File Cohort Contents inside zip ms.zip Multiple Sclerosis ms/reference_annot.h5ad, ms/query_annot.h5ad hp5.zip Human Pancreas hp5/reference.h5ad, hp5/query.h5ad lung.zip Lung-Kim lung/reference_annot.h5ad, lung/query_annot.h5ad cl.zip Cell line cl/reference.h5ad, cl/query.h5ad covid.zip COVID-19 (uncorrected) covid/reference-raw.h5ad, covid/query-raw.h5ad myeloid.zip Myeloid (base) myeloid/reference_adata.h5ad, myeloid/query_adata.h5ad



