five

Hybrid Enterobacteriaceae assemblies using PacBio+Illumina or ONT+Illumina sequencing|基因组组装数据集|测序技术数据集

收藏
Mendeley Data2024-06-25 更新2024-06-27 收录
基因组组装
测序技术
下载链接:
https://figshare.com/articles/dataset/Hybrid_Enterobacteriaceae_assemblies_using_PacBio_Illumina_or_ONT_Illumina_sequencing/7649051/3
下载链接
链接失效反馈
资源简介:
Data associated with: De Maio, Shaw, et al. on behalf of the REHAB consortium (2019), Comparison of long-read sequencing technologies in the hybrid assembly of complex bacterial genomes. biorxiv 530824 Illumina sequencing allows rapid, cheap and accurate whole genome bacterial analyses, but short reads (<300 bp) do not usually enable complete genome assembly. Long read sequencing greatly assists with resolving complex bacterial genomes, particularly when combined with short-read Illumina data (hybrid assembly). However, it is not clear how different long-read sequencing methods impact on assembly accuracy. In this study, we compared hybrid assemblies for 20 bacterial isolates, including two reference strains, using Illumina sequencing and long reads from either Oxford Nanopore Technologies (ONT) or from SMRT Pacific Biosciences (PacBio) sequencing platforms. This set of files includes all hybrid assemblies produced using Unicycler with different sequencing approaches and strategies. Each isolate has 8 hybrid assemblies = 4 x ONT-Illumina + 4 x PacBio-Illumina. There are a total of 158 hybrid assemblies from the full data as two assemblies did not finish (8x20 - 2 = 160 - 2 = 158). Additionally, there are Assemblies were produced from different long read preparation strategies. Hybrid assemblies with Unicycler (n1 = 158): • Basic: no filtering or correction of reads (i.e. all long reads available used for assembly). • Corrected: Long reads were error-corrected and subsampled (preferentially selecting longest reads) to 30-40x coverage using Canu (v1.5, https://github.com/marbl/canu) with default options. • Filtered: long reads were filtered using Filtlong (v0.1.1, https://github.com/rrwick/Filtlong) by using Illumina reads as an external reference for read quality and either removing 10% of the worst reads or by retaining 500Mbp in total, whichever resulted in fewer reads. We also removed reads shorter than 1kb and used the --trim and --split 250 options. • Subsampled: we randomly subsampled long reads to leave approximately 600Mbp (corresponding to a long read coverage around 100x). Long-read only assemblies (n2 = 20 x 2 x 2 = 80):• Flye: we ran Flye (https://github.com/fenderglass/Flye) with the options --plasmids --meta, which have been shown to improve the assemblies of plasmids in bacterial genomes (see: https://github.com/rrwick/Long-read-assembler-comparison) • Pilon: the Flye assemblies were then polished with Illumina short-reads using Pilon (https://github.com/broadinstitute/pilon). Assembly file names have the following format: ${sample-name}_${preparation-strategy}_${long-read-sequencing}.fastae.g. for sample CFT073 the filtered PacBio-Illumina assembly is: CFT073_filtered_pacbio.fasta Also included are assemblies produced after subsampling long-read data to ~10X genome coverage for the following strategies: "basic" (hybrid) and long-read ("flye" and "pilon"). There are n3 = 20 x 3 x 2 = 120 of these assemblies. These have a '10X' preceding the preparation strategy. The total number of assemblies is n1+n2+n3=158+80+120=358. See the associated preprint for more details: https://doi.org/10.1101/530824
创建时间:
2023-06-28
用户留言
有没有相关的论文或文献参考?
这个数据集是基于什么背景创建的?
数据集的作者是谁?
能帮我联系到这个数据集的作者吗?
这个数据集如何下载?
点击留言
数据主题
具身智能
数据集  4099个
机构  8个
大模型
数据集  439个
机构  10个
无人机
数据集  37个
机构  6个
指令微调
数据集  36个
机构  6个
蛋白质结构
数据集  50个
机构  8个
空间智能
数据集  21个
机构  5个
5,000+
优质数据集
54 个
任务类型
进入经典数据集
热门数据集

波士顿房价数据集

波士顿房价数据集是一个经典的机器学习数据集,通常用于回归任务,尤其是房价预测。下方文档中有所有字段顺序的描述。

阿里云天池 收录

ShapeNet

ShapeNet 是由斯坦福大学、普林斯顿大学和美国芝加哥丰田技术研究所的研究人员开发的大型 3D CAD 模型存储库。该存储库包含超过 3 亿个模型,其中 220,000 个模型被分类为使用 WordNet 上位词-下位词关系排列的 3,135 个类。 ShapeNet Parts 子集包含 31,693 个网格,分为 16 个常见对象类(即桌子、椅子、平面等)。每个形状基本事实包含 2-5 个部分(总共 50 个部分类)。

OpenDataLab 收录

中国交通事故深度调查(CIDAS)数据集

交通事故深度调查数据通过采用科学系统方法现场调查中国道路上实际发生交通事故相关的道路环境、道路交通行为、车辆损坏、人员损伤信息,以探究碰撞事故中车损和人伤机理。目前已积累深度调查事故10000余例,单个案例信息包含人、车 、路和环境多维信息组成的3000多个字段。该数据集可作为深入分析中国道路交通事故工况特征,探索事故预防和损伤防护措施的关键数据源,为制定汽车安全法规和标准、完善汽车测评试验规程、

北方大数据交易中心 收录

Allen Brain Atlas

Allen Brain Atlas 是一个综合性的脑图谱数据库,提供了详细的大脑解剖结构、基因表达数据、神经元连接信息等。该数据集包括了小鼠、人类和其他模式生物的大脑数据,旨在帮助研究人员理解大脑的结构和功能。

portal.brain-map.org 收录

yolo-datasets

深度学习目标检测数据集/分割数据集最全最完整的数据集集合,包含电力电气领域、航空影像输电线路与输电塔分割、电力遥感风力发电机、安全带和安全绳检测、变压器漏油故障诊断、高压输电线故障检测、光伏热红外缺陷、风电光伏功率数据、变电站火灾、输电线路语义分割、配网缺陷检测、变电站设备目标检测、太阳能光伏电池板缺陷、pcb电路板检测、绝缘体检测、输电线路防震锤缺陷、电线冰雪覆盖、电力工程电网施工现场安全作业、螺丝识别检测、变电站电力设备的可见光和红外图像、无人机航拍输电线路悬垂线夹、电线线路表面损害、氧化锌避雷器破损识别、热斑光伏发电系统红外热图像等多个领域的数据集。

github 收录