Safe Harbor Canino — processed genome-wide tracks (ATAC-seq, RRBS, Hi-C)
收藏资源简介:
Computational audit snapshot: 25 September 2026. w01 is the priority locus for experimental assessment; bg1k_2848 and bg1k_0442 are proposed genomic comparators. w11 remains conditional: the structural interpretation is unresolved. Seven read groups retain compatible junction geometry in a panel of reported alternatives; this is not whole-genome or independent molecular validation. The wider control pool retains 57 unresolved source ALT entries. No guides or biological safe-harbor validation are claimed. Includes selected source, audit reports and provenance through 2026-09-25. Read RELEASE_20260925.md for scope, exclusions and remaining work. Raw reads, large alignments and downloaded reference inputs are excluded. Existing genome-wide tracks in the dataset record are retained unchanged. Snapshot commit: 607d1bc1e92770d10c98659fc34a283e236d9cb3. Earlier record description (historical) Processed genome-wide tracks from the Safe Harbor Canino project(ROS_Cfam_1.0), companion dataset to the code repository(github.com/stardrako-create/SafeHarborCanine, DOI: 10.5281/zenodo.21996453). - ATAC_*: weighted-mean accessibility, signal variability, and consensus peak frequency across 71 dogs (Jin et al. 2024, PRJNA1048909).- RRBS_*: weighted-mean CpG methylation, coverage frequency, and variability across the same 71 dogs (PRJNA1049514).- HiC_*: multi-resolution contact matrix (.mcool), TAD insulation scores (100/250/500kb windows) and raw insulation table, from Hi-C on a single individual ("Mischka", Wang et al. 2021, PRJNA587469) — intentionally low-weighted structural context, not a population-level track. See the code repository README for full methodology. Methodology — Mother Track construction The ATAC and RRBS tracks are not a naive average across the 71 dogs. Each uses a two-layer weighting scheme: A per-dog global QC weight (ATAC: FRiP + TSS enrichment; RRBS: mapping efficiency + bisulfite conversion rate), min-max normalized across the 71 dogs and rescaled onto [0.2, 1.0] so no dog is ever fully excluded. A per-bin local confidence, comparing each dog's own coverage/depth at that bin against that same dog's genome-wide background rate — this lets under-measured bins drop out of a dog's contribution instead of being averaged in as "closed" or "unmethylated." Consensus ATAC peaks require support from a strict majority of dogs (>=36/71, following ArchR's addReproduciblePeakSet() convention), not a signal-magnitude threshold. TAD boundaries (Hi-C, single individual "Mischka") come from cooltools insulation at 25kb resolution, computed independently at 100kb/250kb/500kb windows, Li-thresholded, boundary called if any window flags it. Full formulas, exact scripts, and parameter values: see METHODS.md (uploaded alongside the data files in this record, and at github.com/stardrako-create/SafeHarborCanine/04_tracks_processadas/ROS_Cfam_1.0/METHODS.md). Update 2026-08-24 — extended 76-dog ATAC cohort (V9-B robustness check) Ehsan Valiollahi sent 5 additional canine ATAC-seq samples (GEO GSE278027, PBMC, 2 breeds — note these trace to the mammary-tumor arm of that study, not healthy controls). Processed identically to the original 71 dogs, then combined two ways: ATAC_ehsan5_*: standalone Mother Track built from just these 5 dogs (same two-layer QC-weight/local-confidence formulas above, normalized to this 5-dog cohort). ATAC_joined76_*: the two Mother Tracks (71-dog + 5-dog) combined by cohort size — joined(bin) = (71*mean_71(bin) + 5*mean_5(bin)) / 76 — not a naive re-pool of 76 equally-weighted individuals. Consensus ATAC peaks recomputed from scratch via bedtools multiinter across all 76 dogs' individual peak calls (threshold rescaled to >=39/76, the same majority-vote rule as the original >=36/71). Rerunning the identical candidate-scoring logic (05_SHIP score_ship_candidates.py) against these ATAC-76 inputs, with every other input held constant, reproduced the same 26/461 surviving candidates and the same top-10 rank order as the 71-dog run, with scores shifting by at most ~0.0006 — a genuine robustness confirmation, not a new result. Full formulas in the updated METHODS.md in this record. Correction 2026-08-24 Two errors in the update above, both caught by Ehsan Valiollahi's review, corrected here: Disease status: the 5 additional dogs (GSE278027, N_172-183) are confirmed healthy controls, not from the mammary-tumor arm as stated above — that was based on a misleading GEO metadata field (a study-wide "tissue" field, not this animal's own status). The source paper's own sample-naming convention (N_=normal, B_=benign, C_=malignant) confirms all 5 are N_ (normal). QC-weight normalization: the 5-dog Mother Track's QC weights were originally min-max normalized relative to just those 5 dogs, which is statistically shaky at n=5 (all 5 actually sit at percentile 45–86 of the true 71-dog FRiP/TSS range, yet the 5-dog-relative normalization gave one of them the floor weight as if it were the worst sample in the project). Fixed by anchoring the normalization to the 71-dog reference range instead. ATAC_ehsan5_* and ATAC_joined76_* files in this version reflect the corrected weights. Neither correction changed the result: rebuilding and rescoring with the corrected weights reproduced the same top-10 candidate list and rank order, scores shifting by at most ~0.0006 total. Full account in METHODS.md (this record) and the SafeHarborCanine repo's 05_SHIP/VERSIONS.md.



