遇见数据集

Improving Gene Annotation of the Peanut Genome by Integrated Proteogenomics Workflow

收藏
NIAID Data Ecosystem2026-03-11 收录
官方服务:

资源简介:

Peanut (Arachis hypogaea L.) is a staple crop in semiarid tropical and subtropical regions. Although the genome of peanut has been fully sequenced, the current gene annotations are still incomplete. New technologies in genomics and proteomics have resulted in the emergence of proteogenomics, which can integrate genomic, transcriptomic, and proteomic data for improving gene annotation. In the present study, we collected RNA-seq and proteomic data from multiple tissues such as seed, shell, and gynophore of peanut and utilized a proteogenomic approach to improve the gene annotation of peanut based on these data. A total of 1 935 655 904 RNA-seq reads and 7 490 280 MS/MS spectra were collected. Ultimately, 13 767 annotated genes were found with evidence at the protein level, and seven novel protein-coding genes were found with both RNA-seq and proteomics evidence. In addition, 35 gene models were updated based on proteomics data. Proteogenomic approaches improved the gene annotation in certain aspects by integrating both RNA-seq and proteomic data. We expect that these approaches could help improve existing genome annotations of other species.

花生(Arachis hypogaea L.)是半干旱热带与亚热带地区的主栽粮食作物。尽管花生基因组已完成全序列测定,但当前的基因注释仍存在不完善之处。基因组学与蛋白质组学领域的新兴技术催生了蛋白质基因组学(proteogenomics),该技术可整合基因组、转录组与蛋白质组数据,以优化基因注释流程。本研究收集了花生种子、种壳、果柄等多种组织的RNA测序(RNA-seq)与蛋白质组数据,并基于此采用蛋白质基因组学方法对花生的基因注释进行优化。本次研究共获取1935655904条RNA-seq读段与7490280条质谱/质谱(MS/MS)谱图。最终,13767个已注释基因获得了蛋白质水平的表达证据,同时发现7个同时具备RNA-seq与蛋白质组学证据的新型蛋白质编码基因。此外,基于蛋白质组学数据更新了35个基因模型。蛋白质基因组学方法通过整合RNA-seq与蛋白质组数据,在多个维度优化了基因注释工作。我们预期该方法可助力其他物种现有基因组注释的完善工作。

创建时间:
2020-05-05
二维码
社区交流群
二维码
科研交流群
商业服务