遇见数据集

False Discovery Versus Familywise Error Rate Approaches to Outlier Detection

收藏
Figshare2016-01-20 更新2026-04-29 收录
官方服务:

资源简介:

Outliers, in general, are observations that deviate sufficiently from a base distribution. This study deals with outlier detection approaches for large samples from continuous univariate distributions. Investigated are the properties of a practical newer outlier detection approach based on use of a false discovery rate (FDR) method in conjunction with a robustly estimated Tukey g-and-h base distribution. Compared are the properties of a boxplot type outlier detection approach that controls the familywise error rate (FWER) with a newer FDR approach. These options are compared in terms of error rates and effects of moving the outliers gradually further from base distribution center while using 5% and 1% FDR or FWERs. Two microarray datasets are used as examples where the assumed null distributions do not fit the data well. In such cases, the proposed estimated Tukey g-and-h null distribution approach leads to superior outlier detection performance. Supplementary materials for this article are available online.

异常值(outlier)通常指与基准分布偏差显著的观测值。本研究针对来自连续单变量分布的大样本,探讨异常值检测方法。本研究考察了一种实用的新型异常值检测方法的特性,该方法将错误发现率(false discovery rate, FDR)方法与稳健估计的图基g-and-h基准分布相结合。本研究对两种方法的特性展开对比:一种是控制家族错误率(familywise error rate, FWER)的箱线图型异常值检测方法,另一种是新型FDR方法。本次对比涵盖错误率指标,以及异常值逐渐偏离基准分布中心所产生的影响,且实验均采用5%和1%的FDR或FWER阈值。本研究使用两个微阵列(microarray)数据集作为示例,此类数据集中假设的零分布(null distribution)与实际数据的拟合效果欠佳。在此类场景下,本文提出的基于稳健估计图基g-and-h零分布的检测方法,可实现更优异的异常值检测性能。本文的补充材料可在线获取。

创建时间:
2016-01-20
二维码
社区交流群
二维码
科研交流群
商业服务