遇见数据集

Material for LORA paper (Veillonella Parvula case)

收藏
Zenodo2025-09-30 更新2026-05-26 收录
官方服务:

资源简介:

Veillonella dataset LORA is a LO Read assembly pipeline part of Sequana (https://github.com/sequana/lora) This dataset corresponds to the Veillonella isolate used in the LORA paper to illustrate bacterial genome assembly from long-read sequencing data. The raw sequencing data originate from the study available at:🔗 PubMed 32817093 Specifically, the dataset is associated with the SRA run: ERR3958992 In the original study, sequencing was performed on a PacBio Sequel II platform. The available files in the SRA project contain only raw subreads in FASTQ format. According to the authors, subreads were used without CCS construction. For the LORA paper, we aimed to generate higher-quality CCS reads using a minimum of 3 passes and RQ ≥ 70%, which required access to the original subreads BAM file. This BAM file was kindly provided by the authors and is shared here to ensure full reproducibility of our analysis. For convenience, we provide the following files: m54091_180306_141024.subreads.fastq.gz — FASTQ file equivalent to the SRA data from the publication. This file contains 338,310 reads. veillonella.subreads.bam — Original PacBio subreads in BAM format, which generated the data referenced in the publication. veillonella.ccs.bam — CCS (Circular Consensus Sequencing) reads generated from the subreads with the following parameters: Minimum read quality (RQ): 0.7 Minimum number of passes: 0These settings produce CCS reads with relatively few passes. veillonella.ccs.fastq.gz — FASTQ version of the above CCS dataset, provided for convenience. These data files enable reproducibility of the analyses presented in the LORA paper and can be reused for independent benchmarking or assembly studies.

提供机构:
Zenodo
创建时间:
2025-09-29
二维码
社区交流群
二维码
科研交流群
商业服务