遇见数据集

Suco - TIL

收藏
Zenodo2026-05-18 更新2026-05-26 收录
官方服务:

资源简介:

Suco (Single cell universal classification omnibus) is a large standardized reference dataset for cell type classification in single cell RNA sequencing data. Suco seeks to tackle the lack of standardized datasets for classification tasks in the single cell genomics field. In other fields of artificial intelligence, like computer vision, standardized datasets such as MNIST or ImageNet have transformed the development of powerful new machine learning methods. Suco features manual uniform standardized hierarchical cell type annotations in independently analyzed datasets. This collection of independent datasets ensures that machine learning classifiers can be tested using statistically independent data & labels. Here, we present the tumor infiltrating leukocyte dataset within Suco which includes > 1000 manual labeled cell type clusters from 8 independent datasets totalling 290 individuals and >0.5 millions cells. ______________________________________________________________________________________________________Updates compared to v1.1 Included cell type hierarchy as nested dictionary in json file. This hierarchy can be used to query the data, e.g. with our Cytopus package wallet-maker/cytopus: Single cell omics biology annotations removed residual low quality cells & doublets increased compression to allow for faster download ______________________________________________________________________________________________________ Structure of the dataset: .zip compressed folder containing datasets from individual studies which have been reprocessed, clustered and annotated independently by two different expert raters (human immunology) Filenames: DATASET_ID.h5ad The dataset ID has the following format (each line followed by ‘-X-‘ separator Tissue/cell type: here PBMC Disease context Publication year First author (optional: followed by _BATCHNAME) DOI (/ in DOI is replace by _ for compatibility with file systems the .h5ad files have the following structure load using the scanpy python package adata = sc.read(FILE_PATH) cell barcode (adata.obs_names) Study-ID + '-X-' + internal barcode adata.obs[‘sample_id’] the sample ID should be the patient ID + '-X-' separator + internal sample ID e.g. TIL-X-BRCA-X-scRNAseq-X-Bassez-X-2021-X-10.1038_s41591-021-01323-8-X-2-X-Pre with TIL-X-BRCA-X-scRNAseq-X-Bassez-X-2021-X-10.1038_s41591-021-01323-8-X-2 being the patient ID -X- the separator adata.obs['patient_id'] dataset id followed by an '-X-' separator and the internal patient id e.g. TIL-X-BRCA-X-scRNAseq-X-Bassez-X-2021-X-10.1038_s41591-021-01323-8-X-35 -X- is the separator 35 is the internal patient adata.obs[cluster_final'] final clustering used for the cell type annotation granularity can differ between subsets --> e.g. clustering from myeloid cells can originate from myeloid subset, clustering from TNK from TNK subset and epithelial from all leukocyte subset should be preceeded by the prefix used for subtyping e.g. 'TNK' for TNK cells followed by a '_' seperator and the cluster number:· e.g. cluster 0 in TNK would be 'TNK_0' cluster 1 in M would be 'M_1' adata.obs[cluster_all'] containing coarse clustering format 'all_CLUSTERNUMBER' adata.obs[‘annotation’] Most granular annotation based on adata.obs[‘cluster_final’] adata.obs[‘annotation_all’] annotation based on adata.obs[‘cluster_all’] ______________________________________________________________________________________________________The dataset contains reprocessed and reannotated data from the following studies. Using this data requires observing the authors' license requirements: 1. Bassez, A. et al. A single-cell map of intratumoral changes during anti-PD1 treatment of patients with breast cancer. Nature Medicine (2021). https://doi.org:10.1038/s41591-021-01323-8 2. Zhang, Y. et al. Single-cell analyses reveal key immune cell subsets associated with response to PD-L1 blockade in triple-negative breast cancer. Cancer Cell 39, 1578-1593.e1578 (2021). https://doi.org:10.1016/j.ccell.2021.09.010 3. Che, L.-H. et al. A single-cell atlas of liver metastases of colorectal cancer reveals reprogramming of the tumor microenvironment in response to preoperative chemotherapy. Cell Discovery 7, 80 (2021). https://doi.org:10.1038/s41421-021-00312-y 4. Moorman, A. et al. Progressive plasticity during colorectal cancer metastasis. Nature 637, 947-954 (2025). https://doi.org:10.1038/s41586-024-08150-0 5. Wang, F. et al. Single-cell and spatial transcriptome analysis reveals the cellular heterogeneity of liver metastatic colorectal cancer. Science Advances 9, eadf5464 (2023). https://doi.org:doi:10.1126/sciadv.adf5464 6. Peng, J. et al. Single-cell RNA-seq highlights intra-tumoral heterogeneity and malignant progression in pancreatic ductal adenocarcinoma. Cell Res 29, 725-738 (2019). https://doi.org:10.1038/s41422-019-0195-y 7. Steele, N. G. et al. Multimodal mapping of the tumor and peripheral blood immune landscape in human pancreatic cancer. Nature Cancer 1, 1097-1112 (2020). https://doi.org:10.1038/s43018-020-00121-4 8. Werba, G. et al. Single-cell RNA sequencing reveals the effects of chemotherapy on human pancreatic adenocarcinoma and its tumor microenvironment. Nature Communications 14, 797 (2023). https://doi.org:10.1038/s41467-023-36296-4

提供机构:
Zenodo
创建时间:
2026-05-18
二维码
社区交流群
二维码
科研交流群
商业服务