Clustering and classification problems in genetics through <i>U</i>-statistics
收藏资源简介:
Genetic data are frequently categorical and have complex dependence structures that are not always well understood. For this reason, clustering and classification based on genetic data, while highly relevant, are challenging statistical problems. Here we consider a versatile <i>U</i>-statistics-based approach for non-parametric clustering that allows for an unconventional way of solving these problems. In this paper we propose a statistical test to assess group homogeneity taking into account multiple testing issues and a clustering algorithm based on dissimilarities within and between groups that highly speeds up the homogeneity test. We also propose a test to verify classification significance of a sample in one of two groups. We present Monte Carlo simulations that evaluate size and power of the proposed tests under different scenarios. Finally, the methodology is applied to three different genetic data sets: global human genetic diversity, breast tumour gene expression and Dengue virus serotypes. These applications showcase this statistical framework's ability to answer diverse biological questions in the high dimension low sample size scenario while adapting to the specificities of the different datatypes.
遗传数据多为分类数据,且往往具备复杂且尚未被完全阐明的依赖结构。因此,基于遗传数据开展聚类与分类研究虽极具现实意义,却属于极具挑战性的统计课题。为此,本文提出一种基于U统计量(U-statistics)的通用非参数聚类方法,为解决此类问题提供了一条非常规路径。本文首先构建了一种兼顾多重检验问题的组同质性评估统计检验,并基于组内与组间相异性提出了一款聚类算法,可大幅提升同质性检验的运算效率;此外,本文还提出了一种用于验证样本在二分类场景下归属类别的分类显著性检验方法。我们开展了蒙特卡洛模拟(Monte Carlo simulation)实验,在多种场景下评估了所提检验的显著性水平与检验功效。最后,将该方法论应用于三类不同的遗传数据集:全球人类遗传多样性数据集、乳腺肿瘤基因表达数据集以及登革病毒(Dengue virus)血清型数据集。上述应用案例表明,该统计框架能够在高维低样本量场景下适配不同数据类型的特性,进而解答多样化的生物学研究问题。




