遇见数据集

hDOM and SGM ref maps for HG002 v1.2 (Nb.BssSI)

收藏
Zenodo2026-09-29 更新2026-10-01 收录
官方服务:

资源简介:

Reference maps for the HG002 v1.2 diploid assembly and the nicking endonuclease Nb.BssSI (CTCGTG), for the two imaging arms of the study. Both are derived from the same assembly and deposited together, separately from the source code, because of their size and because they do not need re-releasing when the code changes. hDOM_reference_map_HG002v12.zip 331 MB fluorescence (hDOM): 47 two-channel TIFF maps at 200 bp/pixel plus genome.maps SGM_reference_maps_HG002v12.zip 3.5 MB SEM (SGM): nick-site fragment maps for maligner_dp The first part of this description covers the hDOM archive, the second the SGM archive. PART 1. hDOM REFERENCE MAP (hDOM_reference_map_HG002v12.zip) CONTENTS 47 TIFF files and one text file, 331 MB compressed: chr01_m, chr01_p ... chr22_m, chr22_p 22 autosomes, both haplotypes chrX_m, chrY_p sex chromosomes (HG002 is male) chrM mitochondrial genome genome.maps in silico Nb.BssSI restriction map Together the TIFFs span 29,997,170 pixels, that is 6.00 Gb of diploid sequence. genome.maps holds 735,853 Nb.BssSI fragments over the same 47 sequences. FILE FORMAT Each TIFF is an 8-bit RGB image, 10 pixels tall (identical rows, for display) and as wide as the chromosome is long at 200 bp per pixel. One pixel is 200 bp. R channel local A/T density, the signal an AT-selective dye reports G channel Nb.BssSI site presence, the signal a labelled nick reports B channel unused, zero Both channels are scaled to 5-255. Red is a histogram of A/T positions smoothed with a Gaussian of sigma = 1 pixel. Green is a histogram of Nb.BssSI sites (CTCGTG and its reverse complement), binarised to presence per pixel, smoothed with the same kernel, then square-rooted before scaling, so that neighbouring sites combine sub-linearly. genome.maps is the site list reduced to inter-site fragment lengths, one line per sequence: name, length in bp, number of fragments, then the fragments. It is the input format used by maligner. SITE POSITIONS ARE AT THE TRUE MOTIF BP genome.maps carries each site at the base pair the motif occurs at, not at the 200 bp pixel it falls in. That is the difference from the HG002 v1.1 record, where every fragment length is a multiple of 200: here 0.5% of fragments are, because a fragment is a difference of two real coordinates. Pinning sites to a multiple of 200 costs accuracy that maligner is sensitive to. Over the 703,252 sites that pair up between HG002 v1.1 and v1.2 the true movement is 6 bp at the median, while a 200 bp lattice records "0 bp" for 95.0% of them and ">=200 bp" for 5.03% with nothing in between, so 4.5% of sites carry an error of about 194 bp that does not exist. maligner consumes the fragment lengths directly, so this shows up in its answer: across the two assembly versions the call is unchanged for 67.0% of molecules on a lattice map against 96.2% on this one. The TIFFs are unaffected. They are 200 bp per pixel either way. HOW THEY WERE BUILT By make_reference.py in hDOM-toolkit, from the HG002 v1.2 chromosome FASTA: python3 make_reference.py --fasta FASTA_DIR --out ref_map --bpp 200 \ --motif CTCGTG --grid snapped is the default and is what produced this record; --grid binned reproduces the v1.1 lattice behaviour. The exact recipe matters. Rebuilding the reference with a nearly correct recipe (red exact, green kernel exact, only the normalisation differing over about 4% of pixels) was enough to move one of seven test molecules to its homologue, which is why the maps are distributed rather than left to be regenerated. HOW TO USE THEM Unzip into a folder and point the pipeline at it: unzip hDOM_reference_map_HG002v12.zip -d ref_map python3 hDOM_pipeline.py INPUT_DIR --ref-map ref_map --bpp 200 The bundled example in hdom/example_genome of the software record searches three molecules against all 47 sequences and checks the result against the recorded answer. Those recorded answers are for this record; the v1.1 maps are a different assembly and place one of the three molecules on the other haplotype. PART 2. SGM REFERENCE MAPS (SGM_reference_maps_HG002v12.zip) CONTENTS Two .maps files and one README, 8 MB uncompressed: genome_smoothed_rm100.maps 789,874 fragments fragments below 100 bp removed genome_unsmoothed.maps 808,027 fragments every predicted site Each file covers 46 sequences, Chr1-22 in both haplotypes (_m and _p) plus ChrX_m and ChrY_p. The unsmoothed map spans 5,999,413,648 bp of diploid sequence. Same format as genome.maps above, with fragment lengths in bp at the true motif position. The two archives' fragment counts differ (808,027 here, 735,853 in genome.maps) by design. The SGM maps list every predicted site. genome.maps lists the peaks of the 200 bp/pixel green channel, where sites within the same pixel or the same smoothing kernel appear as one peak, as they do in a fluorescence image. WHICH FILE TO USE genome_smoothed_rm100.maps is the map the SEM molecules are aligned against. Two nick sites closer than 100 bp, about 34 nm, are a single spot in the SEM image, so the fragment between them is dropped rather than reported as a site the measurement failed to find. Genome-wide this removes 2.2% of fragments. Removing a fragment takes its length with it, so rm100 is 840,030 bp shorter than the unsmoothed map. genome_unsmoothed.maps is the full site list, drawn as the reference track in the alignment figures. Aligning the 53 SEM molecules of the study against genome_smoothed_rm100.maps with maligner_dp at default options reproduces the published alignments byte for byte. FROM HG002 v1.1 The SGM maps of the earlier version were deposited as a separate record (doi:10.5281/zenodo.22703566), which also carried four comparison maps (rm50, rm150, FM_sm200, FM_sm600) that the analysis does not use; they are not rebuilt here. Between the two assemblies the maps have the same fragment counts on every sequence and fragment lengths shift by tens of bp. Across the 53 molecules the best-scoring locus and direction are unchanged, positions move by 52 bp or less, and m-scores by 0.04 or less. PROVENANCE (both archives) Derived from the HG002 v1.2 assembly of the Q100 project (T2T Consortium and Genome in a Bottle), released into the public domain under CC0. The derivation adds only the choices described above and no sequence data beyond what the assembly already contains. RELATED SOFTWARE hDOM-toolkit, which reads both sets of maps: see the software record linked under Related works.

提供机构:
Zenodo
创建时间:
2026-09-29
二维码
社区交流群
二维码
科研交流群
商业服务