遇见数据集

UNC Systems Genetics

收藏
知名数据库2026-06-11 收录
官方服务:

资源简介:

Here we provide genomic sequences for the Collaborative Cross (CC) mouse strains and the eight CC founder strains in the form of FASTA files for the 19 autosomes, sex chromosomes (X and Y), and mitochondria (M). These sequences can be used as reference sequences for high-throughput short-read alignments, or for any other comparative genomic analyses. Each genome comes with a companion MOD file, which can be used to remap coordinates from the FASTA sequences back to reference coordinates. This is necessary since, in general, all gene and genomic annotations are specified relative to the reference. MOD files are genome and version specific, and therefore should always be downloaded together as a set with their associated FASTA sequence. We supply two types of genomes, sequenced and imputed. Sequenced genomes result from direct DNA sequencing at a minimum of 30x coverage, and an iterative alignment process. Imputed genomes are derived from genotype data, where we first construct a haplotype mosaic using MegaMUGA genotypes and then assemble an imputed genome using segments of DNA sequence from the inferred founders

本数据集以FASTA文件格式,提供了协作杂交(Collaborative Cross, CC)小鼠品系及其8个CC创始品系的基因组序列,覆盖19条常染色体、性染色体(X与Y)以及线粒体(M)基因组。上述序列可作为高通量短读长比对的参考序列,或用于各类比较基因组学分析。每个基因组均附带配套的MOD文件,可用于将FASTA序列中的坐标重新映射回原始参考坐标。这一步骤必不可少,因为通常所有基因及基因组注释均基于参考基因组定义。MOD文件具有基因组及版本特异性,因此需与其配套的FASTA序列一同下载使用。本数据集提供两类基因组:测序型与填充型。测序型基因组通过至少30倍覆盖度的直接DNA测序及迭代比对流程获得;填充型基因组则基于基因型数据构建:首先利用MegaMUGA基因型数据构建单倍型嵌合体,再通过来自推断出的创始品系的DNA序列片段组装得到填充型基因组。

搜集汇总
数据集介绍
UNC Systems Genetics 数据集图片
背景与挑战
背景概述
该数据集提供了协作交叉小鼠品系及其八个创始品系的基因组序列,包括FASTA文件和配套的MOD文件,用于高通量短读长比对或其他比较基因组学分析。数据包含序列化和估算两种基因组类型,分别基于直接DNA测序和基因型数据推导。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务