遇见数据集

Randomized chromatin interaction datasets with network structure preserved

收藏
Zenodo2026-01-05 更新2026-05-26 收录
官方服务:

资源简介:

This record provides three randomized chromatin interaction datasets in Hi-C and promoter capture Hi-C (pcHi-C) formats. The randomized datasets were generated using our randomization tool described in Sizovs et al. [1] and implemented in our HiCCliqueGraphs repository [2]. The method produces structurally similar interaction networks by exactly preserving node degree in the network representation (nodes correspond to chromatin segments and edges to interactions), while approximately preserving the distribution of interaction lengths (in base pairs). For each dataset, we provide: the original (unrandomized) interactions, a version with 50% of interactions randomized, a version with 75% of interactions randomized. Source datasets randomized Blood pcHi-C (Javierre et al. [3]). PcHi-C interactions for 17 hematopoietic blood cell types, filtered to retain interactions with count ≥ 5 prior to randomization. Tissue pcHi-C (Jung et al. [4]). PcHi-C interactions for 10 human tissue types, filtered to retain interactions with adjusted p-value ≥ 0.7 prior to randomization. Tissue Hi-C (3DIV / Kim et al. [5]). Hi-C interaction data for 10 human tissue types, using interactions with −log(p-value) > 10. File format All interaction files are provided as CSV (comma-separated) text files without headers.Each row corresponds to a single chromatin interaction between two genomic bins: columns 1–2: start and end coordinates of bin 1 columns 3–4: start and end coordinates of bin 2 Archive structure Files are distributed as ZIP archives with the following directory layout: Top level: select the dataset Second level: select the tissue type or cell type Within each tissue/cell type directory: CSV files are provided per chromosome Acknowledgements The research reported in this study was funded through and supported by the Latvian Council of Scienceproject lzp-2021/1-0236. Citation and references Sizovs, A., et al. A technique for preserving network structure in randomized Hi-C data. Journal of Bioinformatics and Computational Biology, 22(05), 2440001 (2024). IMCS-Bioinformatics. “HiCCliqueGraphs.” GitHub, GitHub, 27 Sept. 2024, https://github.com/IMCS-Bioinformatics/HiCCliqueGraphs Javierre, Biola M., et al. "Lineage-specific genome architecture links enhancers and non-coding disease variants to target gene promoters." Cell 167.5 (2016): 1369-1384. Jung, Inkyung, et al. "A compendium of promoter-centered long-range chromatin interactions in the human genome." Nature genetics 51.10 (2019): 1442-1449. Kim, Kyukwang, et al. "3DIV update for 2021: a comprehensive resource of 3D genome and 3D cancer genome." Nucleic acids research 49.D1 (2021): D38-D46.

本数据集包含三套经随机化处理的染色质交互数据集,格式涵盖Hi-C(高通量染色体构象捕获技术)与启动子捕获Hi-C(promoter capture Hi-C, pcHi-C)。 本套随机化数据集通过我们的随机化工具生成,该工具详见Sizovs等人的研究[1],并已在我们的HiCCliqueGraphs开源仓库[2]中实现。该方法通过严格保留网络表征中的节点度(节点对应染色质片段,边对应染色质交互),生成结构相似的交互网络,同时近似保留交互长度(以碱基对为单位)的分布特征。 针对每套数据集,我们提供以下内容: - 原始(未随机化)交互数据集 - 50%交互经随机化处理的版本 - 75%交互经随机化处理的版本 待随机化的原始数据集: 1. 血液启动子捕获Hi-C(Javierre等人[3]):包含17种造血血细胞类型的pcHi-C交互数据,随机化前已过滤掉交互计数小于5的条目。 2. 组织启动子捕获Hi-C(Jung等人[4]):包含10种人体组织类型的pcHi-C交互数据,随机化前已过滤掉校正后p值小于0.7的条目。 3. 组织Hi-C(3DIV / Kim等人[5]):包含10种人体组织类型的Hi-C交互数据,仅保留−log(p值) >10的交互条目。 文件格式 所有交互文件均以无表头的CSV(逗号分隔值)文本格式提供。每一行对应两个基因组区间之间的单条染色质交互: - 第1-2列:区间1的起始与终止坐标 - 第3-4列:区间2的起始与终止坐标 归档结构 文件以ZIP压缩包形式分发,目录结构如下: - 一级目录:选择目标数据集 - 二级目录:选择组织类型或细胞类型 - 每个组织/细胞类型目录下:按染色体分别提供对应的CSV文件 致谢 本研究相关工作由拉脱维亚科学委员会项目lzp-2021/1-0236资助支持。 引用与参考文献 [1] Sizovs A, 等. 随机化Hi-C数据中保留网络结构的方法. 《生物信息学与计算生物学杂志》, 2024, 22(05): 2440001. [2] IMCS-Bioinformatics. “HiCCliqueGraphs”[EB/OL]. GitHub, 2024年9月27日. https://github.com/IMCS-Bioinformatics/HiCCliqueGraphs [3] Javierre BM, 等. 谱系特异性基因组架构将增强子与非编码疾病变异关联至靶基因启动子. 《细胞》, 2016, 167(5): 1369-1384. [4] Jung I, 等. 人类基因组中以启动子为中心的长程染色质交互汇编. 《自然·遗传学》, 2019, 51(10): 1442-1449. [5] Kim K, 等. 2021年3DIV更新:三维基因组与三维癌症基因组综合资源. 《核酸研究》, 2021, 49(D1): D38-D46.

提供机构:
Zenodo
创建时间:
2026-01-05
二维码
社区交流群
二维码
科研交流群
商业服务