ljybhsw/Genecorpus-104M
收藏资源简介:
Genecorpus-104M是一个大规模预训练语料库,包含约1.04亿个人类单细胞转录组数据,这些数据来自公开可用的广泛组织范围。该语料库用于预训练Geneformer-V2,这是一个预训练的Transformer模型,能够在数据有限的网络生物学设置中进行上下文感知预测。数据集以Huggingface Datasets结构提供标记化数据,基于Apache Arrow格式。每个数据实例代表语料库中单个细胞的秩值编码,这是一种非参数表示方法,通过在每个细胞中按基因表达排名(并基于整个语料库的表达进行归一化)来编码转录组。秩值编码利用语料库中每个基因表达的多次观察,优先区分细胞状态的基因。具体来说,它通过归一化降低普遍高表达的家务基因的排名,同时提高如转录因子等低表达但能高度区分细胞状态的基因在相对上调细胞中的排名。这种方法对可能系统偏差绝对转录计数的技术伪影更具鲁棒性。数据创建动机是映射驱动疾病进展的基因调控网络,以筛选能够通过归一化核心调控元件来纠正网络的分子,而不是靶向可能不改变疾病的外围下游效应器。尽管在罕见疾病或临床难以接近组织等数据有限的情况下,单细胞技术促进了转录组状态的观察,而不需要平均多个细胞的基因表达,为网络交互推断提供了更精确的数据。通过迁移学习概念,利用在大规模通用数据集上预训练的深度学习模型,可以在有限任务特定数据下微调以完成各种下游任务。因此,Genecorpus-104M的组装允许大规模预训练Geneformer,实现在数据有限的网络生物学设置中进行上下文感知预测。
Genecorpus-104M is a large-scale pretraining corpus comprised of approximately 104 million human single cell transcriptomes from a broad range of tissues from publicly available data. This corpus was used for pretraining Geneformer-V2, a pretrained transformer model that enables context-aware predictions in settings with limited data in network biology. The dataset is provided as tokenized data in the Huggingface Datasets structure, based on the Apache Arrow format. Each data instance consists of the rank value encoding for a single cell within the corpus. Rank value encodings provide a nonparametric representation of each single cell’s transcriptome, ranking genes by their expression within that cell normalized by their expression across the entire corpus. This method leverages the many observations of each gene’s expression across the corpus to prioritize genes that distinguish cell state. Specifically, it deprioritizes ubiquitously highly-expressed housekeeping genes by normalizing them to a lower rank, while genes such as transcription factors that may be lowly expressed but highly distinguish cell state are moved to a higher rank in cells where they are relatively upregulated. This rank-based approach is more robust against technical artifacts that may systematically bias absolute transcript counts. The curation rationale is to map gene regulatory networks driving disease progression, enabling screening for molecules that correct the network by normalizing core regulatory elements, rather than targeting peripheral downstream effectors. Although data is limited in settings like rare diseases or clinically inaccessible tissues, single cell technologies facilitate observation of transcriptomic states without averaging gene expression across multiple cells, providing more precise data for network interaction inference. Through transfer learning, deep learning models pretrained on large-scale general datasets can be fine-tuned for downstream tasks with limited task-specific data. Thus, Genecorpus-104M was assembled to allow large-scale pretraining of Geneformer, enabling context-aware predictions in data-limited network biology settings.



