遇见数据集

Reproducible scRNA-seq PBMC Workflow: Reference, Toy Demonstration Data, and Analysis Outputs

收藏
Zenodo2026-03-28 更新2026-05-26 收录
官方服务:

资源简介:

Description This record provides a reproducible, containerized single-cell RNA-seq (scRNA-seq) workflow for peripheral blood mononuclear cell (PBMC) data analysis. Execution is managed via Snakemake and Docker to ensure environment consistency. Workflow Capabilities The integrated pipeline implements an end-to-end orchestration: Upstream: FASTQ QC (FastQC/MultiQC), Optional Trimming (Cutadapt), and STARsolo alignment. Downstream: Seurat-based clustering/annotation, Pseudobulk DESeq2, TOST equivalence testing, GSEA/ORA, and Cell-type–specific co-expression networks (MUUMI). Contents of Version 3.0.0 This version provides a validation suite and a representative result archive (~6 GB): Toy Demonstration Bundle: A chromosome 1 (chr1) mini-reference, pre-built STAR index, and a reduced FASTQ subset (~100,000 reads) for rapid 5–10 minute execution testing. Full Result Archive: Representative execution artifacts from the 4-donor PBMC dataset, including STARsolo gene-cell matrices, Seurat objects, TOST equivalence tables, and MUUMI co-expression networks. The toy dataset validates the orchestration engine through Seurat object creation, while the included results archive provides the technical provenance for the full downstream analysis. Reproducibility Notes The full GRCh38 reference genome and annotation files are not included in this archive due to size. Exact reference sources, versions, and configuration details required to reproduce the full analysis are documented in the associated GitHub repository. The complete workflow, including configuration, container definition, and execution instructions, is available at: https://github.com/Inkasimo/scRNAseq-pbmc-workflow This version reflects a stable snapshot of the demonstration data and representative analysis outputs aligned with the corresponding tagged software release.

描述 本数据集提供一套可复现、容器化的外周血单个核细胞(peripheral blood mononuclear cell, PBMC)单细胞RNA测序(single-cell RNA-seq, scRNA-seq)数据分析工作流。该工作流通过Snakemake与Docker进行任务调度与环境管控,以确保分析环境的一致性。 工作流功能 本集成化流程实现全流程编排: 上游环节:FASTQ质量控制(FastQC/MultiQC)、可选序列修剪(Cutadapt)以及STARsolo序列比对。 下游环节:基于Seurat的细胞聚类与注释、伪bulk差异表达分析(DESeq2)、TOST等效性检验、基因集富集分析(GSEA)/过表达分析(ORA),以及细胞类型特异性共表达网络构建(MUUMI)。 3.0.0版内容 本版本包含一套验证套件与代表性结果归档(大小约6 GB): 演示用小型数据集包:包含1号染色体(chr1)迷你参考基因组、预构建的STAR索引,以及精简后的FASTQ子集(约100,000条reads),可在5至10分钟内完成快速运行测试。 完整结果归档:包含来自4个供体PBMC数据集的代表性运行产物,其中涵盖STARsolo基因-细胞矩阵、Seurat对象、TOST等效性检验表格以及MUUMI共表达网络。 该演示数据集可通过Seurat对象构建验证编排引擎的有效性,而附带的结果归档则为完整下游分析提供了技术溯源依据。 可复现性说明 本归档未包含完整的GRCh38参考基因组与注释文件,因文件体积过大。复现完整分析所需的精确参考来源、版本信息与配置细节,已在关联的GitHub仓库中完成文档化。 完整工作流(含配置、容器定义与运行说明)可通过以下链接获取: https://github.com/Inkasimo/scRNAseq-pbmc-workflow 本版本为演示数据与代表性分析输出的稳定快照,与对应标记的软件发布版本保持对齐。

提供机构:
Zenodo
创建时间:
2026-02-14
二维码
社区交流群
二维码
科研交流群
商业服务