遇见数据集

Material for LORA paper (SGM raw data)

收藏
Zenodo2025-09-30 更新2026-05-26 收录
官方服务:

资源简介:

Streptococcus gallolyticus subsp. macedonicus (SGM) dataset This dataset corresponds to the Streptococcus gallolyticus subsp. macedonicus (SGM) isolate used in the LORA paper to illustrate bacterial genome assembly from long-read sequencing data. The raw sequencing data originate from the study available at: 🔗 PubMed 40810579 The corresponding BioProject is:🔗 PRJNA940176 Specifically, the dataset is associated with the SRA run:🔗 SRR24332397 In the original study, sequencing was performed on a PacBio Sequel II platform. The available files in the SRA project contain only raw subreads in FASTQ format. According to the authors, CCS reads were generated with a minimum of 3 passes and read quality (RQ) ≥ 99%. For the LORA paper, we aimed to generate higher-quality HiFi reads using a minimum of 10 passes and RQ ≥ 99%, which required access to the original subreads BAM file. This BAM file was kindly provided by the authors and is shared here to ensure full reproducibility of our analysis. We therefore provide: The original PacBio subreads BAM file: m54091_201030_155628.subreads.bam (~42 GB) The derived HiFi reads generated with 10 passes and RQ ≥ 99%: hifi.ccs.bam hifi.ccs.fastq.gz (for convenience) These files enable users to reproduce the HiFi read generation and subsequent assembly steps described in the LORA publication.

提供机构:
Zenodo
创建时间:
2025-09-30
二维码
社区交流群
二维码
科研交流群
商业服务