whittle benchmark datasets: length-stratified HG002 ONT subsets
收藏官方服务:
资源简介:
Length-stratified subsets of the public Oxford Nanopore HG002 dorado-SUP release, used to benchmark the whittle long-read trimmer (https://github.com/erdikilic/whittle). Two normalizations are provided: eqbase holds sequence volume constant at 180 Mb per subset (57,883 / 8,848 / 4,103 reads for short / mid / long), and eqread holds read count constant at 8,000 reads per subset. Each of the six subsets spans a distinct read-length regime (N50 approximately 5, 21, and 42 kb) and is provided both as unaligned BAM carrying MM/ML base-modification tags and as gzip-compressed FASTQ. Derived from ONT Open Data.
提供机构:
Zenodo创建时间:
2026-07-14



