遇见数据集

Clustering and classification problems in genetics through <i>U</i>-statistics

收藏
DataCite Commons2020-09-01 更新2024-07-27 收录
官方服务:

资源简介:

Genetic data are frequently categorical and have complex dependence structures that are not always well understood. For this reason, clustering and classification based on genetic data, while highly relevant, are challenging statistical problems. Here we consider a versatile <i>U</i>-statistics-based approach for non-parametric clustering that allows for an unconventional way of solving these problems. In this paper we propose a statistical test to assess group homogeneity taking into account multiple testing issues and a clustering algorithm based on dissimilarities within and between groups that highly speeds up the homogeneity test. We also propose a test to verify classification significance of a sample in one of two groups. We present Monte Carlo simulations that evaluate size and power of the proposed tests under different scenarios. Finally, the methodology is applied to three different genetic data sets: global human genetic diversity, breast tumour gene expression and Dengue virus serotypes. These applications showcase this statistical framework's ability to answer diverse biological questions in the high dimension low sample size scenario while adapting to the specificities of the different datatypes.

遗传数据多为分类数据,且往往具备复杂且尚未被完全阐明的依赖结构。因此,基于遗传数据开展聚类与分类研究虽极具现实意义,却属于极具挑战性的统计课题。为此,本文提出一种基于U统计量(U-statistics)的通用非参数聚类方法,为解决此类问题提供了一条非常规路径。本文首先构建了一种兼顾多重检验问题的组同质性评估统计检验,并基于组内与组间相异性提出了一款聚类算法,可大幅提升同质性检验的运算效率;此外,本文还提出了一种用于验证样本在二分类场景下归属类别的分类显著性检验方法。我们开展了蒙特卡洛模拟(Monte Carlo simulation)实验,在多种场景下评估了所提检验的显著性水平与检验功效。最后,将该方法论应用于三类不同的遗传数据集:全球人类遗传多样性数据集、乳腺肿瘤基因表达数据集以及登革病毒(Dengue virus)血清型数据集。上述应用案例表明,该统计框架能够在高维低样本量场景下适配不同数据类型的特性,进而解答多样化的生物学研究问题。

提供机构:
Taylor & Francis
创建时间:
2017-09-22
搜集汇总
数据集介绍
Clustering and classification problems in genetics through <i>U</i>-statistics 数据集图片
背景与挑战
背景概述
该数据集专注于遗传学中的聚类和分类问题,采用基于U统计的非参数方法,包括组同质性测试、聚类算法和分类显著性测试。数据集通过蒙特卡洛模拟验证方法,并应用于人类遗传多样性、基因表达和病毒血清型等多个实际遗传数据集,适用于高维低样本量场景下的生物问题分析。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务