遇见数据集

Network-Based Segmentation of Biological Multivariate Time Series

收藏
Figshare2016-01-18 更新2026-04-29 收录
官方服务:

资源简介:

Molecular phenotyping technologies (e.g., transcriptomics, proteomics, and metabolomics) offer the possibility to simultaneously obtain multivariate time series (MTS) data from different levels of information processing and metabolic conversions in biological systems. As a result, MTS data capture the dynamics of biochemical processes and components whose couplings may involve different scales and exhibit temporal changes. Therefore, it is important to develop methods for determining the time segments in MTS data, which may correspond to critical biochemical events reflected in the coupling of the system’s components. Here we provide a novel network-based formalization of the MTS segmentation problem based on temporal dependencies and the covariance structure of the data. We demonstrate that the problem of partitioning MTS data into segments to maximize a distance function, operating on polynomially computable network properties, often used in analysis of biological network, can be efficiently solved. To enable biological interpretation, we also propose a breakpoint-penalty (BP-penalty) formulation for determining MTS segmentation which combines a distance function with the number/length of segments. Our empirical analyses of synthetic benchmark data as well as time-resolved transcriptomics data from the metabolic and cell cycles of Saccharomyces cerevisiae demonstrate that the proposed method accurately infers the phases in the temporal compartmentalization of biological processes. In addition, through comparison on the same data sets, we show that the results from the proposed formalization of the MTS segmentation problem match biological knowledge and provide more rigorous statistical support in comparison to the contending state-of-the-art methods.

分子表型技术(Molecular phenotyping technologies,如转录组学(transcriptomics)、蛋白质组学(proteomics)与代谢组学(metabolomics))可同步获取生物系统内不同层级信息处理与代谢转化环节所产生的多变量时间序列(Multivariate Time Series, MTS)数据。由此,多变量时间序列数据能够捕捉生化过程及其组分的动态变化,这些过程与组分的耦合往往涉及不同尺度,并随时间动态演化。因此,开发用于识别多变量时间序列数据中时间分段的方法至关重要——这些时间分段往往对应着系统组分耦合所反映的关键生化事件。本研究提出了一种基于网络、依托时间依赖性与数据协方差结构的全新多变量时间序列分割问题形式化框架。我们证明,将多变量时间序列数据划分为若干分段以最大化某一基于生物网络分析中常用的多项式可计算网络属性构建的距离函数的问题,可被高效求解。为便于生物学解读,我们还提出了断点惩罚(breakpoint-penalty, BP-penalty)公式化方法用于多变量时间序列分割,该方法将距离函数与分段数量/长度相结合。我们通过对合成基准数据集以及酿酒酵母(Saccharomyces cerevisiae)代谢周期与细胞周期的时间分辨转录组学数据开展实证分析,证明所提方法能够准确推断生物过程时序分区的各个阶段。此外,通过在相同数据集上的对比实验,我们发现相较于当前同类顶尖前沿方法,本研究提出的多变量时间序列分割问题形式化方法所得结果更契合生物学认知,且具备更严谨的统计学支撑。

创建时间:
2016-01-18
二维码
社区交流群
二维码
科研交流群
商业服务