遇见数据集

Genome assemblies of Drosophila melanogaster inbred strains from Hemker et al. 2026

收藏
Figshare2026-03-03 更新2026-04-28 收录
官方服务:

资源简介:

The increasing accessibility of long-read sequencing and the rapid development of automated variant callers are promoting the generation of population-level structural variation data. However, the effect of the length of long-reads on automated variant callers is not well understood, especially for non-human species. We used Oxford Nanopore Technologies to long-read sequence eight, inbred *D. melanogaster *strains to extremely high coverage (mean 238x), and we then downsampled the reads (to 30x-coverage) to create read pools of different length distributions. Here are the assembled genomes from each of these read-length distributions. Average distributions stats are described in the table below.Description of the data and file structureEight inbred D. melanogaster strains were deeply sequenced (mean: 238x) with nanopore long reads. These read pools were then computationally downsampled into five pools of distinct read-length distributions and coverages. Each of these pools were then assembled into genomes. In total there are five assemblies per strain and 40 assemblies in total.The assemblies are provided in a tarballs, grouped by read-length distribution.Files and variablesEach set of eight assemblies for a given read-length distribution is included in the tarball named [distribution]_assemblies.tar.gz.Download and access these assemblies with tar xzf [distribution]_assemblies.tar.gzEach tarball will have the following contents:[strain_id].[distribution].rm.fastaWhere [strain_id] is one of dmel11, dmel12, dmel19, dmel21, dmel22, dmel31, dmel37, dmel55.Strain-specific information can be found in the Materials and Methods and Supplementary Information

长读长测序技术的可及性日益提升,自动化变异检测工具(automated variant callers)也迎来快速发展,有力推动了群体水平结构变异数据的产出。然而,长读长序列的长度对自动化变异检测工具的影响机制尚未得到充分阐释,针对非人类物种的相关研究尤为匮乏。 本研究采用牛津纳米孔科技(Oxford Nanopore Technologies)平台,对8个近交系黑腹果蝇(*D. melanogaster*)品系进行超高覆盖度长读长测序,平均测序深度达238倍;随后通过计算方法将测序reads下采样至30倍覆盖度,构建得到不同长度分布的reads混合样本。本数据集包含上述各长度分布reads对应的组装基因组,各长度分布的平均统计信息详见下表。 数据与文件结构说明:本研究对8个近交系黑腹果蝇品系进行了深度纳米孔长读长测序,平均测序深度为238倍。随后通过计算手段将reads混合样本下采样为5组具有不同长度分布和测序覆盖度的样本,并分别对每组样本进行基因组组装。每个品系对应5套组装结果,总计40套组装基因组。 所有组装基因组按照reads长度分布进行分组,并打包为tar.gz压缩包。 文件与变量说明:针对某一特定reads长度分布的8套组装基因组,将被打包至名为`[distribution]_assemblies.tar.gz`的压缩包中。可通过命令`tar xzf [distribution]_assemblies.tar.gz`下载并解压获取组装结果。每个压缩包包含以下文件:`[strain_id].[distribution].rm.fasta`,其中`[strain_id]`为以下8个品系ID之一:dmel11、dmel12、dmel19、dmel21、dmel22、dmel31、dmel37、dmel55。各品系的详细信息可参阅研究的材料与方法部分及补充材料。

创建时间:
2026-03-03
二维码
社区交流群
二维码
科研交流群
商业服务