five

Hybrid Enterobacteriaceae assemblies using PacBio+Illumina or ONT+Illumina sequencing|基因组组装数据集|测序技术数据集

收藏
Mendeley Data2024-06-25 更新2024-06-27 收录
基因组组装
测序技术
下载链接:
https://figshare.com/articles/dataset/Hybrid_Enterobacteriaceae_assemblies_using_PacBio_Illumina_or_ONT_Illumina_sequencing/7649051/3
下载链接
链接失效反馈
资源简介:
Data associated with: De Maio, Shaw, et al. on behalf of the REHAB consortium (2019), Comparison of long-read sequencing technologies in the hybrid assembly of complex bacterial genomes. biorxiv 530824 Illumina sequencing allows rapid, cheap and accurate whole genome bacterial analyses, but short reads (<300 bp) do not usually enable complete genome assembly. Long read sequencing greatly assists with resolving complex bacterial genomes, particularly when combined with short-read Illumina data (hybrid assembly). However, it is not clear how different long-read sequencing methods impact on assembly accuracy. In this study, we compared hybrid assemblies for 20 bacterial isolates, including two reference strains, using Illumina sequencing and long reads from either Oxford Nanopore Technologies (ONT) or from SMRT Pacific Biosciences (PacBio) sequencing platforms. This set of files includes all hybrid assemblies produced using Unicycler with different sequencing approaches and strategies. Each isolate has 8 hybrid assemblies = 4 x ONT-Illumina + 4 x PacBio-Illumina. There are a total of 158 hybrid assemblies from the full data as two assemblies did not finish (8x20 - 2 = 160 - 2 = 158). Additionally, there are Assemblies were produced from different long read preparation strategies. Hybrid assemblies with Unicycler (n1 = 158): • Basic: no filtering or correction of reads (i.e. all long reads available used for assembly). • Corrected: Long reads were error-corrected and subsampled (preferentially selecting longest reads) to 30-40x coverage using Canu (v1.5, https://github.com/marbl/canu) with default options. • Filtered: long reads were filtered using Filtlong (v0.1.1, https://github.com/rrwick/Filtlong) by using Illumina reads as an external reference for read quality and either removing 10% of the worst reads or by retaining 500Mbp in total, whichever resulted in fewer reads. We also removed reads shorter than 1kb and used the --trim and --split 250 options. • Subsampled: we randomly subsampled long reads to leave approximately 600Mbp (corresponding to a long read coverage around 100x). Long-read only assemblies (n2 = 20 x 2 x 2 = 80):• Flye: we ran Flye (https://github.com/fenderglass/Flye) with the options --plasmids --meta, which have been shown to improve the assemblies of plasmids in bacterial genomes (see: https://github.com/rrwick/Long-read-assembler-comparison) • Pilon: the Flye assemblies were then polished with Illumina short-reads using Pilon (https://github.com/broadinstitute/pilon). Assembly file names have the following format: ${sample-name}_${preparation-strategy}_${long-read-sequencing}.fastae.g. for sample CFT073 the filtered PacBio-Illumina assembly is: CFT073_filtered_pacbio.fasta Also included are assemblies produced after subsampling long-read data to ~10X genome coverage for the following strategies: "basic" (hybrid) and long-read ("flye" and "pilon"). There are n3 = 20 x 3 x 2 = 120 of these assemblies. These have a '10X' preceding the preparation strategy. The total number of assemblies is n1+n2+n3=158+80+120=358. See the associated preprint for more details: https://doi.org/10.1101/530824
创建时间:
2023-06-28
用户留言
有没有相关的论文或文献参考?
这个数据集是基于什么背景创建的?
数据集的作者是谁?
能帮我联系到这个数据集的作者吗?
这个数据集如何下载?
点击留言
数据主题
具身智能
数据集  4099个
机构  8个
大模型
数据集  439个
机构  10个
无人机
数据集  37个
机构  6个
指令微调
数据集  36个
机构  6个
蛋白质结构
数据集  50个
机构  8个
空间智能
数据集  21个
机构  5个
5,000+
优质数据集
54 个
任务类型
进入经典数据集
热门数据集

OpenSonarDatasets

OpenSonarDatasets是一个致力于整合开放源代码声纳数据集的仓库,旨在为水下研究和开发提供便利。该仓库鼓励研究人员扩展当前的数据集集合,以增加开放源代码声纳数据集的可见性,并提供一个更容易查找和比较数据集的方式。

github 收录

China Health and Nutrition Survey (CHNS)

China Health and Nutrition Survey(CHNS)是一项由美国北卡罗来纳大学人口中心与中国疾病预防控制中心营养与健康所合作开展的长期开放性队列研究项目,旨在评估国家和地方政府的健康、营养与家庭计划政策对人群健康和营养状况的影响,以及社会经济转型对居民健康行为和健康结果的作用。该调查覆盖中国15个省份和直辖市的约7200户家庭、超过30000名个体,采用多阶段随机抽样方法,收集了家庭、个体以及社区层面的详细数据,包括饮食、健康、经济和社会因素等信息。自2011年起,CHNS不断扩展,新增多个城市和省份,并持续完善纵向数据链接,为研究中国社会经济变化与健康营养的动态关系提供了重要的数据支持。

www.cpc.unc.edu 收录

PCLT20K

PCLT20K数据集是由湖南大学等机构创建的一个大规模PET-CT肺癌肿瘤分割数据集,包含来自605名患者的21,930对PET-CT图像,所有图像都带有高质量的像素级肿瘤区域标注。该数据集旨在促进医学图像分割研究,特别是在PET-CT图像中肺癌肿瘤的分割任务。

arXiv 收录

CE-CSL

CE-CSL数据集是由哈尔滨工程大学智能科学与工程学院创建的中文连续手语数据集,旨在解决现有数据集在复杂环境下的局限性。该数据集包含5,988个从日常生活场景中收集的连续手语视频片段,涵盖超过70种不同的复杂背景,确保了数据集的代表性和泛化能力。数据集的创建过程严格遵循实际应用导向,通过收集大量真实场景下的手语视频材料,覆盖了广泛的情境变化和环境复杂性。CE-CSL数据集主要应用于连续手语识别领域,旨在提高手语识别技术在复杂环境中的准确性和效率,促进聋人与听人社区之间的无障碍沟通。

arXiv 收录

RDD2022

RDD2022是一个多国图像数据集,用于自动道路损伤检测,由印度理工学院罗凯里分校交通系统中心等机构创建。该数据集包含来自六个国家的47,420张道路图像,标注了超过55,000个道路损伤实例。数据集通过智能手机和高分辨率相机等设备采集,旨在通过深度学习方法自动检测和分类道路损伤。RDD2022数据集的应用领域包括道路状况的自动监测和计算机视觉算法的性能基准测试,特别关注于解决多国道路损伤检测的问题。

arXiv 收录