遇见数据集

Suco - PBMC

收藏
Zenodo2025-05-07 更新2026-05-26 收录
官方服务:

资源简介:

Suco (Single cell universal classification omnibus) is a large standardized reference dataset for cell type classification in single cell RNA sequencing data. Suco seeks to tackle the lack of standardized datasets for classification tasks in the single cell genomics field. In other fields of artificial intelligence, like computer vision, standardized datasets such as MNIST or ImageNet have transformed the development of powerful new machine learning methods. Suco features manual uniform standardized hierarchical cell type annotations in independently analyzed datasets. This collection of independent datasets ensures that machine learning classifiers can be tested using statistically independent data & labels. Here, we present the peripheral blood mononuclear cell dataset within Suco which includes > 1200 independent manual cell type cluster labels from 12 datasets totalling >500 individuals and >5 millions cells. ______________________________________________________________________________________________________ Structure of the dataset: .zip compressed folder containing datasets from individual studies which have been reprocessed, clustered and annotated independently by two different expert raters (human immunology) Filenames: DATASET_ID.h5ad The dataset ID has the following format (each line followed by ‘-X-‘ separator Tissue/cell type: here PBMC Disease context Publication year First author (optional: followed by _BATCHNAME) DOI (/ in DOI is replace by _ for compatibility with file systems the .h5ad files have the following structure load using the scanpy python package adata = sc.read(FILE_PATH) cell barcode (adata.obs_names) Study-ID + '-X-' + internal barcode adata.obs[‘sample_id’] the sample ID should be the patient ID + '-X-' separator + internal sample ID e.g. TIL-X-BRCA-X-scRNAseq-X-Bassez-X-2021-X-10.1038_s41591-021-01323-8-X-2-X-Pre with TIL-X-BRCA-X-scRNAseq-X-Bassez-X-2021-X-10.1038_s41591-021-01323-8-X-2 being the patient ID -X- the separator adata.obs['patient_id'] dataset id followed by an '-X-' separator and the internal patient id e.g. TIL-X-BRCA-X-scRNAseq-X-Bassez-X-2021-X-10.1038_s41591-021-01323-8-X-35 -X- is the separator 35 is the internal patient adata.obs[cluster_final'] final clustering used for the cell type annotation granularity can differ between subsets --> e.g. clustering from myeloid cells can originate from myeloid subset, clustering from TNK from TNK subset and epithelial from all leukocyte subset should be preceeded by the prefix used for subtyping e.g. 'TNK' for TNK cells followed by a '_' seperator and the cluster number:· e.g. cluster 0 in TNK would be 'TNK_0' cluster 1 in M would be 'M_1' adata.obs[cluster_all'] containing coarse clustering format 'all_CLUSTERNUMBER' adata.obs[‘annotation’] Most granular annotation based on adata.obs[‘cluster_final’] adata.obs[‘annotation_all’] annotation based on adata.obs[‘cluster_all’]

提供机构:
Zenodo
创建时间:
2025-03-30
二维码
社区交流群
二维码
科研交流群
商业服务