UNC Systems Genetics
收藏资源简介:
Here we provide genomic sequences for the Collaborative Cross (CC) mouse strains and the eight CC founder strains in the form of FASTA files for the 19 autosomes, sex chromosomes (X and Y), and mitochondria (M). These sequences can be used as reference sequences for high-throughput short-read alignments, or for any other comparative genomic analyses. Each genome comes with a companion MOD file, which can be used to remap coordinates from the FASTA sequences back to reference coordinates. This is necessary since, in general, all gene and genomic annotations are specified relative to the reference. MOD files are genome and version specific, and therefore should always be downloaded together as a set with their associated FASTA sequence. We supply two types of genomes, sequenced and imputed. Sequenced genomes result from direct DNA sequencing at a minimum of 30x coverage, and an iterative alignment process. Imputed genomes are derived from genotype data, where we first construct a haplotype mosaic using MegaMUGA genotypes and then assemble an imputed genome using segments of DNA sequence from the inferred founders
本数据集以FASTA文件格式,提供了协作杂交(Collaborative Cross, CC)小鼠品系及其8个CC创始品系的基因组序列,覆盖19条常染色体、性染色体(X与Y)以及线粒体(M)基因组。上述序列可作为高通量短读长比对的参考序列,或用于各类比较基因组学分析。每个基因组均附带配套的MOD文件,可用于将FASTA序列中的坐标重新映射回原始参考坐标。这一步骤必不可少,因为通常所有基因及基因组注释均基于参考基因组定义。MOD文件具有基因组及版本特异性,因此需与其配套的FASTA序列一同下载使用。本数据集提供两类基因组:测序型与填充型。测序型基因组通过至少30倍覆盖度的直接DNA测序及迭代比对流程获得;填充型基因组则基于基因型数据构建:首先利用MegaMUGA基因型数据构建单倍型嵌合体,再通过来自推断出的创始品系的DNA序列片段组装得到填充型基因组。




