Maize (Zea mays L.) Genome Diversity as Revealed by RNA-Sequencing
收藏资源简介:
Maize is rich in genetic and phenotypic diversity. Understanding the sequence, structural, and expression variation that contributes to phenotypic diversity would facilitate more efficient varietal improvement. RNA based sequencing (RNA-seq) is a powerful approach for transcriptional analysis, assessing sequence variation, and identifying novel transcript sequences, particularly in large, complex, repetitive genomes such as maize. In this study, we sequenced RNA from whole seedlings of 21 maize inbred lines representing diverse North American and exotic germplasm. Single nucleotide polymorphism (SNP) detection identified 351,710 polymorphic loci distributed throughout the genome covering 22,830 annotated genes. Tight clustering of two distinct heterotic groups and exotic lines was evident using these SNPs as genetic markers. Transcript abundance analysis revealed minimal variation in the total number of genes expressed across these 21 lines (57.1% to 66.0%). However, the transcribed gene set among the 21 lines varied, with 48.7% expressed in all of the lines, 27.9% expressed in one to 20 lines, and 23.4% expressed in none of the lines. De novo assembly of RNA-seq reads that did not map to the reference B73 genome sequence revealed 1,321 high confidence novel transcripts, of which, 564 loci were present in all 21 lines, including B73, and 757 loci were restricted to a subset of the lines. RT-PCR validation demonstrated 87.5% concordance with the computational prediction of these expressed novel transcripts. Intriguingly, 145 of the novel de novo assembled loci were present in lines from only one of the two heterotic groups consistent with the hypothesis that, in addition to sequence polymorphisms and transcript abundance, transcript presence/absence variation is present and, thereby, may be a mechanism contributing to the genetic basis of heterosis.
玉米具有丰富的遗传与表型多样性。解析与表型多样性相关的序列、结构及表达变异,可助力更高效的品种改良。RNA测序(RNA-seq)是开展转录组分析、评估序列变异并鉴定新型转录本序列的高效手段,在玉米这类大型、复杂且富含重复序列的基因组中应用价值尤为突出。本研究对21个玉米自交系的全幼苗RNA进行了测序,这些自交系涵盖了多样的北美及外来种质资源。通过单核苷酸多态性(SNP)检测,共鉴定出351710个分布于全基因组的多态性位点,覆盖22830个已注释基因。以这些SNP作为遗传标记进行聚类分析,可观察到两个独立的杂种优势群与外来种质类群各自形成紧密的聚类簇。转录本丰度分析显示,21个品系的表达基因总数占比差异极小,区间为57.1%至66.0%。但21个品系间的转录基因集合存在差异:其中48.7%的基因在所有品系中均有表达,27.9%的基因仅在1至20个品系中表达,另有23.4%的基因在所有品系中均未表达。对未比对至参考B73基因组序列的RNA-seq读长进行从头组装(de novo assembly),共获得1321个高可信度的新型转录本。其中564个位点在包括B73在内的全部21个品系中均存在,757个位点仅存在于部分品系中。经逆转录聚合酶链式反应(RT-PCR)验证,其与这些新型表达转录本的计算预测结果的吻合度达87.5%。值得关注的是,145个从头组装的新型位点仅存在于两个杂种优势群中的其中一个类群,这支持如下假说:除序列多态性与转录本丰度外,转录本存在/缺失变异同样广泛存在,其或为构成杂种优势遗传基础的重要机制之一。



