Data from: Whole genome sequencing of elite rice cultivars as a comprehensive information resource for marker assisted selection
收藏资源简介:
Current advances in sequencing technologies and bioinformatics revealed the genomic background of rice, a staple food for the poor people, and provided the basis to develop large genomic variation databases for thousands of cultivars. Proper analysis of this massive resource is expected to give novel insights into the structure, function, and evolution of the rice genome, and to aid the development of rice varieties through marker assisted selection or genomic selection. In this work we present sequencing and bioinformatics analyses of 104 rice varieties belonging to the major subspecies of Oryza sativa. We identified repetitive elements and recurrent copy number variation covering about 200 Mbp of the rice genome. Genotyping of over 18 million polymorphic locations within O. sativa allowed us to reconstruct the individual haplotype patterns shaping the genomic background of elite varieties used by farmers throughout the Americas. Based on a reconstruction of the alleles for the gene GBSSI, we could identify novel genetic markers for selection of varieties with high amylose content. We expect that both the analysis methods and the genomic information described here would be of great use for the rice research community and for other groups carrying on similar sequencing efforts in other crops.
当前,测序技术与生物信息学领域的最新进展揭示了作为贫苦民众主粮的水稻的基因组背景,并为构建覆盖数千份栽培品种的大型基因组变异数据库奠定了基础。对这一海量资源开展合理分析,有望为解析水稻基因组的结构、功能与演化提供全新见解,并借助标记辅助选择或基因组选择助力水稻品种的培育。本研究针对隶属于栽培稻(Oryza sativa)主要亚种的104份水稻品种,开展了测序与生物信息学分析。我们鉴定出覆盖水稻基因组约200 Mbp的重复序列元件与频发拷贝数变异。对栽培稻内超过1800万个多态位点开展基因分型后,我们得以重构出塑造美洲各地农户所使用的优良品种基因组背景的单倍型模式。基于对GBSSI基因等位变异的重构,我们鉴定出可用于选育高直链淀粉含量品种的新型遗传标记。我们认为,本文所述的分析方法与基因组信息,将对水稻研究学界以及其他在其他作物中开展同类测序研究的团队具有重要应用价值。



