遇见数据集

Pangenome Graph Datasets for Snarl and Superbubble Detection Experiments

收藏
Zenodo2025-11-20 更新2026-05-26 收录
官方服务:

资源简介:

This record contains cleaned blunt GFA graphs used in the experimental evaluation of BubbleFinder, a method for detecting snarls and superbubbles using BC–SPQR tree decomposition. All graphs were processed with a Snakemake workflow and correspond to standardized versions of pangenome, variation, and de Bruijn graphs. Each file is a gzip-compressed blunt GFA (GFA1) with the suffix .cleaned.gfa.gz. Included graphs: Pangenome graphs (PGGB-based) ecoli50.cleaned.gfa.gz: E. coli pangenome graph from the PGGB "ecoli50" dataset. primates_chr6.cleaned.gfa.gz: primate chromosome 6 pangenome graph. mouse17_chr19.cleaned.gfa.gz: mouse chromosome 19 pangenome graph (PGGB "mouse17" dataset). tomato23_chr2.cleaned.gfa.gz: tomato chromosome 2 pangenome graph (PGGB "tomato23" dataset). Additional E. coli pangenome graph coli3682.cleaned.gfa.gz: large E. coli pangenome graph ("coli3682" dataset). Human variation graphs (1000 Genomes Project, GRCh37) Ultrabubble_dataset_chr1.cleaned.gfa.gz Ultrabubble_dataset_chr10.cleaned.gfa.gz Ultrabubble_dataset_chr22.cleaned.gfa.gz These human variation graphs were constructed from Phase 3 VCFs on GRCh37 using vg construct and converted to GFA with vg convert. They correspond to the "ultrabubble dataset" used for method comparison. De Bruijn pangenome graph GCA.cleaned.gfa.gz: order-41 de Bruijn graph built from 10 Myxococcus xanthus assemblies using GGCAT. Processing Raw GFAs (from PGGB, VG, GGCAT, or public releases) were first normalized into blunt graphs using GetBlunted, then cleaned by removing non-essential header lines to produce the .cleaned.gfa files. These cleaned blunt GFAs are exactly the graph inputs used for all tools in our benchmarks (BubbleFinder, BubbleGun, and vg snarls).

提供机构:
Zenodo
创建时间:
2025-11-20
二维码
社区交流群
二维码
科研交流群
商业服务