Signal recovery in single cell batch integration
收藏资源简介:
Data integration to align cells across batches has become a cornerstone of single cell data analysis, critically affecting downstream results. Yet, how much biological signal is erased during integration? Currently, there are no guidelines for when the biological differences between samples are separable from batch effects, and thus, data integration involves a lot of guesswork: Cells across batches should be aligned to be appropriately mixed, while preserving main cell type clusters. We show evidence that current paradigms for single cell data integration are unnecessarily aggressive, removing biologically meaningful variation and introducing distortion. To remedy this, we present a novel statistical model and computationally scalable algorithm, CellANOVA, that harnesses experimental design to explicitly recover biological signals that are erased during single cell data integration. CellANOVA utilizes a pool-of-controls design concept, applicable across diverse settings, to separate unwanted variation from biological variation of interest. When applied with existing integration methods, CellANOVA allows the recovery of subtle biological signals and corrects, to a large extent, the data distortion introduced by integration. Further, CellANOVA explicitly estimates cell- and gene-specific batch effect terms which can be used to identify the cell types and pathways exhibiting the largest batch variations, providing clarity as to which biological signals can be recovered. These concepts are illustrated on studies of diverse designs, where the biological signals that are recovered by CellANOVA are validated by orthogonal assays. In particular, we show that CellANOVA is effective in the challenging case of single-cell and single-nuclei data integration, where it recovered subtle biological signals are can be validated and replicated by external data.
跨批次对齐细胞的数据集成现已成为单细胞数据分析的核心基石,其对下游分析结果具有决定性影响。然而,在数据集成过程中究竟会抹去多少生物学信号?目前尚无针对「样本间生物学差异可与批次效应(batch effects)」分离场景的指导准则,因此数据集成往往依赖大量经验推测:跨批次的细胞需在对齐后实现合理混合,同时保留主要细胞类型聚类结果。本研究证实,当前主流的单细胞数据集成范式存在过度激进的问题,会移除具有生物学意义的变异并引入数据失真。为解决这一问题,我们提出了一款全新的统计模型与计算可扩展算法CellANOVA,该方法利用实验设计原理,可精准恢复单细胞数据集成过程中被抹去的生物学信号。CellANOVA采用「对照池(pool-of-controls)」设计理念,可适配多种应用场景,能够将非目标变异与目标生物学变异分离开来。当与现有集成方法联用时,CellANOVA可恢复微弱的生物学信号,并在极大程度上校正集成过程引入的数据失真。此外,CellANOVA可精准估算细胞与基因特异性的批次效应项,借此可识别出批次变异程度最高的细胞类型与生物学通路,从而明确哪些生物学信号可被成功恢复。我们通过多种实验设计的研究案例阐释了上述理念,其中CellANOVA恢复的生物学信号均通过正交实验(orthogonal assays)得到了验证。尤为重要的是,本研究证实CellANOVA在单细胞与单细胞核数据集成这一高难度场景中效果显著,其恢复的微弱生物学信号可通过外部数据得到验证与复现。



