uce-spleen-traditional-vs-uce-comparison
收藏资源简介:
该数据集是一个用于比较 UCE(Universal Cell Embedding)与传统 Scanpy 注释方法在小鼠脾脏单细胞 RNA 测序(scRNA-seq)数据上性能的比对包。数据来源于 Rad54b 外显子3 敲除(KO)与野生型(WT)小鼠脾脏,每组 3 个生物学重复,均为雄性 C57BL/6 品系。数据集包含约 20 MB 的压缩包,内含比较图表(PNG 和 PDF 格式)、混淆矩阵、细胞组成 log2 差异倍数(log2FC)、Leiden 聚类标记物以及汇总 JSON 文件。此外,还提供了传统标记物面板(CSV 长格式/宽格式及 JSON),这些标记物是手动整理的经典小鼠脾脏谱系基因,而非从该数据集差异表达中自动选取。传统注释流程为:标准化 → 选择高变基因(2000个)→ 缩放 → PCA → Harmony(按样本校正)→ Leiden 聚类(分辨率0.8)→ 基于簇的 z 分数标记物 argmax 分配。关键结果显示:广泛标签一致性约为 70.5%,精细标签一致性约为 60.8%(合并 T 亚型后为 68.6%);粒细胞在 KO 组中的扩增趋势在两种方法中一致(UCE log2FC +2.66,传统方法 +2.54);主要差异在于 UCE 过度识别 NK 细胞(约55%的 UCE NK 细胞被传统方法归类为 B 细胞)。该数据集仅包含图表和表格,不提供原始计数矩阵或 h5ad 文件。适用于评估单细胞注释方法的准确性、一致性及生物学差异挖掘。许可协议为 CC BY 4.0。
This dataset is a comparison package for evaluating the performance of UCE (Universal Cell Embedding) versus traditional Scanpy annotation methods on mouse spleen single-cell RNA sequencing (scRNA-seq) data. The data originates from Rad54b exon 3 knockout (KO) and wild-type (WT) mouse spleens, with three biological replicates per group, all from male C57BL/6 strain. The dataset includes an approximately 20 MB compressed archive containing comparison charts (PNG and PDF formats), confusion matrices, log2 fold changes (log2FC) of cell composition, Leiden cluster markers, and summary JSON files. Additionally, it provides traditional marker panels (CSV long/wide formats and JSON), which are manually curated classical mouse spleen lineage genes, not automatically selected from differential expression of this dataset. The traditional annotation pipeline is: normalization → selection of highly variable genes (2000) → scaling → PCA → Harmony (batch correction by sample) → Leiden clustering (resolution 0.8) → cluster-based z-score marker argmax assignment. Key results show: broad label consistency ~70.5%, fine label consistency ~60.8% (68.6% after merging T subtypes); granulocyte expansion in KO group is consistent between methods (UCE log2FC +2.66, traditional +2.54); major discrepancy is UCE over-identification of NK cells (~55% of UCE NK cells classified as B cells by traditional method). The dataset contains only charts and tables, not raw count matrices or h5ad files. It is suitable for evaluating accuracy, consistency, and biological difference discovery of single-cell annotation methods. License: CC BY 4.0.
小鼠脾脏 scRNA-seq:UCE精修与传统Scanpy注释对比数据集
数据集概览
该数据集提供了一份Rad54b外显子3敲除(KO)与野生型(WT)小鼠脾脏单细胞RNA测序数据的对比分析包,每组包含3个生物学重复(雄性C57BL/6小鼠)。核心目标是比较UCE(Universal Cell Embedding)与传统Scanpy流程在细胞类型注释上的差异。
数据内容
| 路径 | 描述 |
|---|---|
UCE_vs_Traditional_comparison_bundle.zip |
完整可下载压缩包(约20 MB) |
figures/ |
对比图(PNG和PDF格式) |
tables/ |
混淆矩阵、组成log2FC、Leiden标记、汇总JSON文件 |
markers/ |
传统标记基因面板(CSV长/宽格式及JSON) |
传统分析流程
标准化 → 高变基因(2000) → 缩放 → PCA → Harmony批次校正 → Leiden聚类(分辨率=0.8) → 基于簇级z-score标记基因argmax注释
标记基因为人工整理的经典小鼠脾脏谱系基因(见markers/),并非从本数据集的差异表达中自动筛选。
关键结果
- 宽标签一致性:约70.5%
- 细标签一致性:约60.8%(T亚型合并后为约68.6%)
- 粒细胞KO扩增一致性:UCE log2FC +2.66 vs 传统 +2.54
- 主要分歧:UCE过度识别NK细胞(约55%的UCE NK细胞在传统注释中为B细胞)
详细结果见tables/comparison_summary.json和figures/FigCompare_main_panel.png。
方法与引用说明
- UCE方法来源:Rosen等,Nature(Universal Cell Embedding)
- 传统流程:Scanpy标准工作流 + Harmony批次校正
- 注意:本数据集仅包含图和表,不含原始计数矩阵或h5ad文件
许可协议
CC BY 4.0





