遇见数据集

Human Skin Cell Atlas data objects

收藏
Zenodo2026-06-30 更新2026-08-13 收录
官方服务:

资源简介:

Human Skin Cell Atlas This upload contains processed data objects (.h5ad files) generated for the Human Skin Cell Atlas (HSCA) project and associated publication. NOTE! The most up-to-date per-cell metadata is provided in the Metadata_HSCA_CellMetadata_Core_Extended.csv data table. Individual .h5ad files may contain metadata entries that differ from those in Metadata_HSCA_CellMetadata_Core_Extended.csv. In cases where metadata values differ between files, the information in Metadata_HSCA_CellMetadata_Core_Extended.csv should be considered the authoritative version. Metadata_HSCA_CellMetadata_Core_Extended.csv contains per-cell metadata for both the core and extended HSCA datasets, including: Study and experiment information. Donor (including bodysite annotation), sample, library preparation, sequencing, and mapping metadata. Quality-control metrics. Annotations: Original author annotations Integrated HSCA annotation levels 1–4 for core HSCA cells. Level 1 annotation for extended HSCA cells. The deposited datasets include: Separate AnnData objects for different annotation levels and cell-type subsets. Mapping and integration results with external datasets. An extended HSCA dataset object. Spatial transcriptomics mapping results for Stereo-seq and Xenium datasets. Metadata tables, cluster annotation, color mapping, and differential gene expression analysis files. Filename prefixes and suffixes: Metadata: files containing metadata regarding cells, samples or used cluster colors HSCA: Human Skin Cell Atlas core or extended datasets. Depending on the full filename, these files may represent the core HSCA atlas, a specific HSCA cell-type subset, a subcluster-level object, or the extended HSCA dataset. Steele: Data derived from the Steele et al. dataset after integration with, or mapping onto, the HSCA reference. Forsthuber: Data derived from the Forsthuber et al. cancer-associated fibroblast (CAF) dataset after integration with, or mapping onto, the HSCA reference. ObRe: Data derived from the Ober-Reynolds et al. dataset after integration with, or mapping onto, the HSCA reference. ObRe is used as a shorthand for Ober-Reynolds. Stereoseq: Stereo-seq data with Stereoscope-based mapping to the HSCA reference. Corresponding nuclei images for the included samples are provided in Stereoseq_nuclei_images.zip. Xenium: Xenium data with cell2location-based mapping to the HSCA reference. Corresponding images are included within the uploaded Xenium .zarr archive. Additional files: Metadata_Cluster_annotation_and_colors.py contains cluster-number mappings, full cluster names, abbreviated cluster names, and associated cluster colors. DGEA_L1_L2_L3_L4_Subclusters.zip contains differential gene expression analysis results for all subclustering levels, from level 1 to level 4. Metadata_SupplementaryTable_SampleMetadata_Extended.xlsx contains per-sample metadata for the core and extended HSCA dataset, including study information, donor identifiers, basic donor metadata such as sex and age, and anatomical-region annotations. Metadata_HSCA_CellMetadata_Core_Extended.csv contains the most up-to-date per-cell metadata for the core and extended HSCA datasets. This includes study, experiment, donor, sample, library preparation, sequencing, mapping, anatomical site, quality-control, original annotation, and integrated HSCA annotation metadata. README_general_annotation_metadata.md contains detailed information about metadata column contents and .h5ad data structure Differential gene expression analysis: For each analysis group, each subcluster was compared against all other cells within the same group. For example, a level 2 keratinocyte cluster was compared against all other keratinocytes. For each level and cell type, the differential expression results are provided in three formats: .csv files for programmatic access. .parquet files for efficient programmatic access. .xlsx files for manual browsing and exploration, with each subcluster provided as a separate sheet. Gene filtering: All AnnData objects were filtered to retain a curated set of standard genes obtained from BioMart. The retained genes include protein-coding and non-coding genes, while spike-ins, technical controls, and non-standard constructs were removed. This filtering resulted in approximately 35,000–40,000 genes per dataset. AnnData structure: Raw counts are stored in .X as a sparse matrix. Log-normalized counts are stored in adata.layers as a sparse matrix.

提供机构:
Zenodo
创建时间:
2026-06-30
二维码
社区交流群
二维码
科研交流群
商业服务