遇见数据集

Identification of Replication Timing Domains Using DNN-HMM. Homo sapiens

收藏
NIAID Data Ecosystem2026-03-08 收录
官方服务:

资源简介:

Purpose: Sixteen GSM Samples from GSE34399 was used to identify four different types of replication timing domains. Methods: 1. Chromosome 1 of Bj_Rep1 was manually annotated. 2. We developed a new supervised method called DNN-HMM (Deep Neural Network-Hidden Markov Model), and used the manual annotation as the training set to learn a model. 3. The model learnt in Step 2 was used to divide the un-annotated Repli-seq datas into four different replication domains (early replication domain, down transition zone, late replication domain, up transition zone). Result: We used DNN-HMM to identify four different replication timing domains respectively in fifteen cell lines. The accuracy of identification was about 87%, and the overlapping percentage of two independent replicates of Bj cell line was about 83%. Data File Formats :(bed) chrom - The name of the chromosome chromStart - The starting position of the feature in the chromosome chromEnd - The ending position of the feature in the chromosome category - The domain identified (denoted by: ERD, short for early replication domain; DTZ, short for down transition zone; LRD, short for late replication domain; UTZ, short for up transition zone) Overall design: Repli-seq datas from GSE34399 were used as the raw datas. After mapping and normalization, the signals from six cell cycle fractions: G1/G1b, S1, S2, S3, S4, G2 (six fraction profile) were merged into a matrix. The six-dimensional matrix then was used as input of DNN-HMM to identify replication timing domains.

研究目的:本研究采用GSE34399数据集下的16个GSM样本,旨在识别四类不同的复制时序结构域。 方法:1. 对Bj_Rep1的1号染色体进行人工注释;2. 开发一种新型监督学习方法——深度神经网络-隐马尔可夫模型(Deep Neural Network-Hidden Markov Model,缩写DNN-HMM),并以人工注释结果作为训练集完成模型训练;3. 利用步骤2训练得到的模型,将未注释的Repli-seq数据划分为四类复制结构域:早期复制结构域(early replication domain,缩写ERD)、下降过渡区(down transition zone,缩写DTZ)、晚期复制结构域(late replication domain,缩写LRD)以及上升过渡区(up transition zone,缩写UTZ)。 结果:我们使用DNN-HMM在15个细胞系中分别识别出四类复制时序结构域,模型识别准确率约为87%;Bj细胞系的两组独立生物学重复的重叠率约为83%。 数据文件格式:采用BED格式,各字段说明如下: - chrom:染色体名称 - chromStart:染色体上特征片段的起始位置 - chromEnd:染色体上特征片段的终止位置 - category:所识别的结构域(缩写对应关系:ERD为早期复制结构域(early replication domain);DTZ为下降过渡区(down transition zone);LRD为晚期复制结构域(late replication domain);UTZ为上升过渡区(up transition zone)) 实验设计:以GSE34399数据集的Repli-seq数据作为原始数据,经比对与标准化处理后,将G1/G1b、S1、S2、S3、S4、G2共6个细胞周期组分的信号整合为矩阵,随后将该六维矩阵作为DNN-HMM的输入,以完成复制时序结构域的识别。

创建时间:
2014-01-10
搜集汇总
数据集介绍
Identification of Replication Timing Domains Using DNN-HMM. Homo sapiens 数据集图片
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务