遇见数据集

Summary of Processed Multi-Omics and GSEA Results in KLF2-Deficient and KLF2-Overexpressing Cells

收藏
Mendeley Data2026-05-21 收录
官方服务:

资源简介:

For RNA-seq analyses, sequencing reads were aligned to the GRCm38.p6 mouse reference genome using STAR. Gene-level read counts (raw counts) were summarized using featureCounts. Differential expression analysis was conducted with DESeq2, which uses a negative binomial distribution for statistical modeling. The Benjamini-Hochberg method was applied to adjust p-values for multiple comparisons. Gene set enrichment analysis (GSEA) and functional gene set enrichment were conducted using the gseapy, utilizing the C7 Immunological Signatures collections from the Molecular Signatures Database. For ATAC-seq, Activated P14 cells transduced with empty vector- or Klf2-expressing vector were cultured and expanded in vitro for 2.5 days. cells were sorted for ATAC-seq. In brief, 2×150-bp paired-end reads obtained from NovaSeq were trimmed for Nextera adaptors using cutadapt (v4.8, with parameters -m 36 -n 3 -q 10) and aligned to the mouse genome (GRCm38.p6) using Bowtie2 (v2.5.1, with parameters --very-sensitive -X 2000 --no-mixed --no-discordant). Duplicate reads were marked using Picard (v2.9.5, https://broadinstitute.github.io/picard/), and only non-duplicated, properly paired reads were retained using samtools (v1.18, with parameters --F 1804 -f 2 --q 20). Reads were adjusted for Tn5 transposase shifts (+4 bp for the sense strand and −5 bp for the antisense strand), pooled by sample type, and peaks were called using MACS2 (v2.2.7.1, with parameters -q 0.05 --nolambda --keep-dup all --call-summits -f BAMPE). Peaks were finalized by retaining those with higher cut-offs (MACS2 -q 0.05) and were subsequently consolidated into reproducible consensus peaks across replicates (≥50% reproducibility) using Diffbind (v2.10.0, https://bioconductor.org/packages/release/bioc/html/DiffBind.html ). Differentially accessible (DA) peaks were generated using dba.analyze function (bFullLibrarySize = FALSE ) in Diffbind. Gene set enrichment analysis (GSEA) and functional gene set enrichment were conducted using the gseapy, utilizing the C7 Immunological Signatures collections from the Molecular Signatures Database.

在RNA测序(RNA-seq)分析中,使用STAR工具将测序读段比对至GRCm38.p6小鼠参考基因组。利用featureCounts对基因水平的读段计数(原始计数)进行汇总。采用DESeq2进行差异表达分析,该工具以负二项分布开展统计建模。通过Benjamini-Hochberg法校正多重比较的P值。使用gseapy进行基因集富集分析(Gene Set Enrichment Analysis, GSEA)及功能基因集富集分析,数据取自分子特征数据库(Molecular Signatures Database)的C7免疫特征集(C7 Immunological Signatures collections)。 转座酶可及性测序(ATAC-seq)部分:将转导空载体或过表达Klf2载体的活化P14细胞体外培养并扩增2.5天后,分选细胞用于ATAC-seq实验。简言之,从NovaSeq测序平台获得的2×150 bp双端读段,使用cutadapt(v4.8,参数为-m 36 -n 3 -q 10)去除Nextera接头序列,随后利用Bowtie2(v2.5.1,参数为--very-sensitive -X 2000 --no-mixed --no-discordant)比对至小鼠基因组GRCm38.p6。使用Picard(v2.9.5,https://broadinstitute.github.io/picard/)标记重复读段,再通过samtools(v1.18,参数为--F 1804 -f 2 --q 20)仅保留非重复且正确配对的读段。对读段进行Tn5转座酶偏移校正(正义链+4 bp,反义链-5 bp),按样本类型合并读段后,使用MACS2(v2.2.7.1,参数为-q 0.05 --nolambda --keep-dup all --call-summits -f BAMPE)进行峰调用。最终保留MACS2校正阈值(-q 0.05)下的峰,随后通过Diffbind(v2.10.0,https://bioconductor.org/packages/release/bioc/html/DiffBind.html)将重复样本间可重复(≥50%重复率)的峰整合为一致性可及峰。使用Diffbind中的dba.analyze函数(参数bFullLibrarySize = FALSE)生成差异可及(DA)峰。 再次使用gseapy进行基因集富集分析(GSEA)及功能基因集富集分析,数据取自分子特征数据库的C7免疫特征集。

创建时间:
2026-04-22
二维码
社区交流群
二维码
科研交流群
商业服务