Randomized chromatin interaction datasets with network structure preserved
收藏资源简介:
This record provides three randomized chromatin interaction datasets in Hi-C and promoter capture Hi-C (pcHi-C) formats. The randomized datasets were generated using our randomization tool described in Sizovs et al. [1] and implemented in our HiCCliqueGraphs repository [2]. The method produces structurally similar interaction networks by exactly preserving node degree in the network representation (nodes correspond to chromatin segments and edges to interactions), while approximately preserving the distribution of interaction lengths (in base pairs). For each dataset, we provide: the original (unrandomized) interactions, a version with 50% of interactions randomized, a version with 75% of interactions randomized. Source datasets randomized Blood pcHi-C (Javierre et al. [3]). PcHi-C interactions for 17 hematopoietic blood cell types, filtered to retain interactions with count ≥ 5 prior to randomization. Tissue pcHi-C (Jung et al. [4]). PcHi-C interactions for 10 human tissue types, filtered to retain interactions with adjusted p-value ≥ 0.7 prior to randomization. Tissue Hi-C (3DIV / Kim et al. [5]). Hi-C interaction data for 10 human tissue types, using interactions with −log(p-value) > 10. File format All interaction files are provided as CSV (comma-separated) text files without headers.Each row corresponds to a single chromatin interaction between two genomic bins: columns 1–2: start and end coordinates of bin 1 columns 3–4: start and end coordinates of bin 2 Archive structure Files are distributed as ZIP archives with the following directory layout: Top level: select the dataset Second level: select the tissue type or cell type Within each tissue/cell type directory: CSV files are provided per chromosome Acknowledgements The research reported in this study was funded through and supported by the Latvian Council of Scienceproject lzp-2021/1-0236. Citation and references Sizovs, A., et al. A technique for preserving network structure in randomized Hi-C data. Journal of Bioinformatics and Computational Biology, 22(05), 2440001 (2024). IMCS-Bioinformatics. “HiCCliqueGraphs.” GitHub, GitHub, 27 Sept. 2024, https://github.com/IMCS-Bioinformatics/HiCCliqueGraphs Javierre, Biola M., et al. "Lineage-specific genome architecture links enhancers and non-coding disease variants to target gene promoters." Cell 167.5 (2016): 1369-1384. Jung, Inkyung, et al. "A compendium of promoter-centered long-range chromatin interactions in the human genome." Nature genetics 51.10 (2019): 1442-1449. Kim, Kyukwang, et al. "3DIV update for 2021: a comprehensive resource of 3D genome and 3D cancer genome." Nucleic acids research 49.D1 (2021): D38-D46.



