Is it time to stop sweeping data cleaning under the carpet? A novel algorithm for outlier management in growth data
收藏资源简介:
All data are prone to error and require data cleaning prior to analysis. An important example is longitudinal growth data, for which there are no universally agreed standard methods for identifying and removing implausible values and many existing methods have limitations that restrict their usage across different domains. A decision-making algorithm that modified or deleted growth measurements based on a combination of pre-defined cut-offs and logic rules was designed. Five data cleaning methods for growth were tested with and without the addition of the algorithm and applied to five different longitudinal growth datasets: four uncleaned canine weight or height datasets and one pre-cleaned human weight dataset with randomly simulated errors. Prior to the addition of the algorithm, data cleaning based on non-linear mixed effects models was the most effective in all datasets and had on average a minimum of 26.00% higher sensitivity and 0.12% higher specificity than other methods. Data cleaning methods using the algorithm had improved data preservation and were capable of correcting simulated errors according to the gold standard; returning a value to its original state prior to error simulation. The algorithm improved the performance of all data cleaning methods and increased the average sensitivity and specificity of the non-linear mixed effects model method by 7.68% and 0.42% respectively. Using non-linear mixed effects models combined with the algorithm to clean data allows individual growth trajectories to vary from the population by using repeated longitudinal measurements, identifies consecutive errors or those within the first data entry, avoids the requirement for a minimum number of data entries, preserves data where possible by correcting errors rather than deleting them and removes duplications intelligently. This algorithm is broadly applicable to data cleaning anthropometric data in different mammalian species and could be adapted for use in a range of other domains.
所有数据均易出现误差,因此在开展分析前需进行数据清洗(data cleaning)。其中一个典型示例为纵向生长数据(longitudinal growth data):目前尚无公认的标准方法可用于识别并剔除不合理数值,且现有多数方法存在局限性,无法跨领域推广应用。本研究设计了一种基于预定义截断值(cut-off)与逻辑规则相结合的决策算法,可对生长测量数据进行修正或删除。针对5种不同的纵向生长数据集,本研究分别测试了添加与未添加该算法的5种生长数据清洗方法:其中4组为未经过清洗的犬只体重或身高数据集,1组为带有随机模拟误差的已预清洗人类体重数据集。在未添加该算法时,基于非线性混合效应模型(non-linear mixed effects models)的数据清洗方法在所有数据集上均表现最优,其灵敏度平均较其他方法高出至少26.00%,特异度平均高出0.12%。使用该算法的数据清洗方法可提升数据留存率,并能够按照金标准(gold standard)修正模拟误差,将数值恢复至误差模拟前的原始状态。该算法可提升所有数据清洗方法的性能,其中非线性混合效应模型方法的平均灵敏度与特异度分别提升了7.68%与0.42%。将非线性混合效应模型与该算法结合用于数据清洗,可通过重复的纵向测量数据允许个体生长轨迹(growth trajectories)与群体轨迹存在差异,能够识别连续误差或首次数据录入时出现的误差,无需满足最小数据录入量要求,尽可能通过修正而非删除的方式留存数据,并可智能去除重复数据。该算法可广泛应用于不同哺乳类物种的人体测量学数据(anthropometric data)清洗工作,且可适配至诸多其他领域。




