遇见数据集

Control T-cell receptor (TCR) alpha and beta chain nucleotide and amino acid sequences from human and mouse

收藏
Zenodo2026-04-11 更新2026-05-26 收录
官方服务:

资源简介:

A dataset of pooled T-cell receptor (TCR) sequences for TCR alpha and beta chains of human and mouse. Sequences are obtained from various samples of healthy individuals/mice using our conventional protocols: see for example [Britanova et al "Dynamics of individual T cell repertoires: from cord blood to centenarians" The Journal of Immunology 2016] and [Izraelson et al. "Comparative analysis of murine T‐cell receptor repertoires." Immunology 2018]. The sequences are stored as gzipped clonotype tables in VDJtools format, see [https://vdjtools-doc.readthedocs.io/en/master/input.html#vdjtools-format]. This control dataset can be used as a proxy for a generative VDJ rearrangement model to estimate the expected frequency distribution of TCRs and check for enrichment of rare TCR clonotypes and groups of similar TCR sequences. For the implementation of the enrichment analysis, please see CalcDegreeStats routine from VDJtools software, see [https://vdjtools-doc.readthedocs.io/en/master/annotate.html#calcdegreestats]. Files named "human.tra.strict.txt.gz", etc are pools of random/naive TCR clonotypes containing unique V/J/CDR3 nucleotide sequence combinations observed in data. The pools.zip file is used for TCR motif inference in VDJdb database [https://github.com/antigenomics/vdjdb-motifs], it contains human.tra.aa.txt, etc files that contain random/naive TCR clonotypes grouped by CDR3 amino acid sequence with the most frequent representative V and J. Additional files (2026 update) We provide precomputed TCR embeddings and background datasets generated using the TCREmP method. The archive `redcea_bg.gz` contains background repertoires for HomoSapiens both TRA and TRB chains, including: - `*_background_100k.tsv` — background clonotypes in AIRR format - `*_background_embeddings.parquet` — vector representations (TCREmP embeddings) of TCR sequences These embeddings are designed for downstream analysis of TCR similarity, clustering, and enrichment detection. In particular, they can be used as a reference background for methods that identify antigen-associated TCRs or enriched clusters of similar receptors. The embeddings were generated using the TCREmP pipeline, which encodes TCR sequences (CDR3 + VJ gene information) into a continuous vector space preserving sequence similarity. For methodological details, see:https://www.sciencedirect.com/science/article/abs/pii/S0022283625002712 This dataset complements the original clonotype pools by providing a ready-to-use representation space for machine learning and statistical analysis of immune repertoires. The provided background embeddings can be used as a proxy distribution for repertoire-level statistical testing and cluster-based enrichment analysis.

提供机构:
Zenodo
创建时间:
2026-04-11
二维码
社区交流群
二维码
科研交流群
商业服务