OmniCellTOSG
收藏资源简介:
OmniCellTOSG数据集是由华盛顿大学圣路易斯分校的研究团队创建的,该数据集整合了来自不同组织、疾病和细胞类型的1200万单个细胞的单细胞转录组数据。通过收集CellxGene、GEO、Brain Cell Atlas和SEA-AD等多个来源的数据,经过严格的质量控制和标准化预处理,形成了包含547,168个细胞的最终数据集。每个细胞或元细胞都与标签(如器官、疾病、性别、年龄、细胞亚型)相关联的信号图/系统,旨在通过图推理解码细胞信号系统。该数据集为解码复杂的细胞信号系统提供了新的图数据模型,并促进了大规模预训练语言模型和图神经网络模型的开发。
The OmniCellTOSG dataset was constructed by a research team from Washington University in St. Louis. This dataset aggregates single-cell transcriptomic data from 12 million single cells spanning various tissues, diseases and cell types. By collecting data from multiple repositories including CellxGene, GEO, Brain Cell Atlas and SEA-AD, and applying rigorous quality control and standardized preprocessing workflows to the acquired data, the team generated a final dataset containing 547,168 cells. Each cell or metacell is associated with signal graphs or systems annotated with labels such as organ, disease, sex, age and cell subtype, and this dataset is designed to decode cellular signaling systems via graph reasoning. This dataset provides a novel graph-based data model for decoding complex cellular signaling systems, and facilitates the development of large-scale pre-trained language models and graph neural network models.

- 1OmniCellTOSG: The First Cell Text-Omic Signaling Graphs Dataset for Joint LLM and GNN Modeling华盛顿大学圣路易斯分校 · 2025年



