遇见数据集

Segmenting accelerometer data from daily life with unsupervised machine learning

收藏
Figshare2019-01-09 更新2026-04-29 收录
官方服务:

资源简介:

PurposeAccelerometers are increasingly used to obtain valuable descriptors of physical activity for health research. The cut-points approach to segment accelerometer data is widely used in physical activity research but requires resource expensive calibration studies and does not make it easy to explore the information that can be gained for a variety of raw data metrics. To address these limitations, we present a data-driven approach for segmenting and clustering the accelerometer data using unsupervised machine learning.MethodsThe data used came from five hundred fourteen-year-old participants from the Millennium cohort study who wore an accelerometer (GENEActiv) on their wrist on one weekday and one weekend day. A Hidden Semi-Markov Model (HSMM), configured to identify a maximum of ten behavioral states from five second averaged acceleration with and without addition of x, y, and z-angles, was used for segmenting and clustering of the data. A cut-points approach was used as comparison.ResultsTime spent in behavioral states with or without angle metrics constituted eight and five principal components to reach 95% explained variance, respectively; in comparison four components were identified with the cut-points approach. In the HSMM with acceleration and angle as input, the distributions for acceleration in the states showed similar groupings as the cut-points categories, while more variety was seen in the distribution of angles.ConclusionOur unsupervised classification approach learns a construct of human behavior based on the data it observes, without the need for resource expensive calibration studies, has the ability to combine multiple data metrics, and offers a higher dimensional description of physical behavior. States are interpretable from the distributions of observations and by their duration.

研究目的与背景:加速度计正愈发广泛地被用于获取健康研究中具有重要价值的身体活动特征数据。切点法(cut-points approach)是身体活动研究中分割加速度计数据的常用方法,但该方法需要依赖高资源消耗的校准研究,且难以灵活探索各类原始数据指标可提供的信息。为克服上述局限性,本研究提出一种数据驱动的方法,借助无监督机器学习(unsupervised machine learning)实现加速度计数据的分割与聚类。研究方法:本研究使用的数据来自千禧队列研究(Millennium cohort study)中的514名14岁参与者,这些参与者在1个工作日与1个休息日的手腕处佩戴了GENEActiv加速度计。本研究采用隐半马尔可夫模型(Hidden Semi-Markov Model, HSMM),该模型被配置为可从5秒平均加速度数据(加入或未加入x、y、z轴角度信息)中识别最多10种行为状态,用于数据的分割与聚类。同时采用切点法作为对照方法。研究结果:当输入特征包含或不包含角度指标时,为达到95%的解释方差(explained variance),行为状态分类分别需要8个和5个主成分(principal components);相较之下,切点法仅需4个主成分即可达成该目标。在以加速度与角度作为输入的HSMM模型中,各行为状态下的加速度分布与切点法的分类组别具有相似的聚类特征,而角度分布则展现出更多样的变化。研究结论:我们提出的无监督分类方法无需开展高资源消耗的校准研究,即可基于观测数据构建人类行为的表征模型,同时能够整合多种数据指标,并为身体行为提供更高维度的描述。各行为状态可通过观测数据的分布及其持续时长得到合理的解释。

创建时间:
2019-01-09
二维码
社区交流群
二维码
科研交流群
商业服务