遇见数据集

Supplementary Data for "MetaTCR: A Framework for Analyzing Batch Effects in TCR Repertoire Datasets"

收藏
Zenodo2026-08-18 更新2026-08-20 收录
官方服务:

资源简介:

Overview MetaTCR is a computational framework for standardizing T-cell receptor (TCR) repertoires and mitigating batch effects in adaptive immune receptor repertoire sequencing (AIRR-seq) data. It constructs a population-scale reference TCR space and projects individual repertoires onto it, turning variable-length repertoires into fixed-dimensional feature profiles (meta-vectors) that enable robust cross-study comparison, dataset integration, and batch-effect correction. This record provides the large data files needed to reproduce the study and to reuse the reference space on new data. Analysis code, the small reference artifacts (primary cluster centroids and functional-cluster mappings), and per-sample metadata are maintained in the MetaTCR GitHub repository (https://github.com/deepomicslab/MetaTCR). Please cite this record together with that repository. Contents The deposit is split into four compressed archives; each extracts to a single top-level folder that mirrors the paths expected by the MetaTCR pipeline. 1. database.tar.gz — reference database and embeddings Note: to use MetaTCR only for encoding your own repertoires, you do not need this archive — ready-to-use reference cluster centroids are already provided in the code repository. MetaTCR_reference_fullseq.txt — the frozen reference set of TCRβ clonotypes, each composed of a CDR3 with its V and J gene (5,329,945 representative clonotypes assembled and deduplicated across many studies); the input used to build the reference TCR space. reference_embeddings/ — pre-computed TCR2vec embeddings of the reference clonotypes: 22 sharded NumPy arrays (shard_*.npy, float16) that concatenate, in shard order, to a (5,329,945 × 120) matrix aligned 1-to-1 with the sequence file. Includes reference_embedding_manifest.tsv (per-shard index + SHA-256 checksums) and a README.md describing the format and how to read the shards. 2. repertoire_data.tar.gz — processed input repertoires Per-sample TCRβ repertoire tables (TSV), one directory per study, for the 19 cohorts used in the downstream analyses and figures. Each repertoire is quality-controlled and reduced to its high-frequency, representative clonotypes (each defined by CDR3 + V/J gene usage). These are the raw inputs consumed by the MetaTCR encoding step. 3. encoding.tar.gz — MetaTCR meta-vectors MetaTCR-encoded feature matrices (Python pickle, .pk) for the downstream datasets — the fixed-dimensional meta-vectors summarizing each sample's abundance and diversity profiles over the 1,024 reference clusters. These are the encoded representations used in the benchmarking, distance-metric, and integration experiments. 4. antigen.tar.gz — antigen-specificity ground truth TCR–epitope annotations from the McPAS-TCR database, used to validate that the functional clusters concentrate epitope-matched TCRs. McPAS-TCR_filt_ept_full_deduplicated.tsv — the filtered, deduplicated McPAS-TCR reference set. McPAS_with_primary_clusters.csv — a balanced evaluation subset with each TCR assigned to its reference cluster (the direct input to the antigen-specificity validation). Usage Extract the archives so that database/, repertoire_data/, and encoding/ sit under the project data/ tree, then follow the step-by-step scripts in the GitHub repository (reference construction, repertoire encoding, distance/metric computation, and batch integration).

提供机构:
Zenodo
创建时间:
2026-08-18
二维码
社区交流群
二维码
科研交流群
商业服务