遇见数据集

Murine datasets and computational analysis: Prdm16_pos antigen-presenting cells

收藏
Zenodo2025-04-14 更新2026-05-26 收录
官方服务:

资源简介:

Prdm16-dependent antigen-presenting cells induce tolerance to gut antigens This repository contains all code and computational analysis for the murine datasets associated with the above paper. Items are organized as follows: 2024-04-26/ contains Delta+7kb mouse model raw sequencing data. 2024-06-21/ contains Rorc(t)_DeltaCD11c mouse model raw sequencing data. 2024-10-16/ contains Multiome (RNA+ATAC) raw sequencing data. 2022AkagbosuEtAl/ contains the Multiome public dataset from PMID 36070798 that is integrated alongside our similar data. (Generated by downloading from GEO repository GSE174405 and initial processing through CellRanger with default parameters). 2014-11-26/ contains bulk RNA sequencing data in bigWig format, as input for the Broad IGV tool, to ultimately generate Figure 1b. Delta7kB_Murine_Analysis/ contains all code in R to process finalized datasets, ultimately generating Figures 2a-c, and ExtData Fig 4a-e. Cd11c_Murine_Analysis/ contains all code in R to process finalized datasets, ultimately generating Figure 2d. Multiome_Murine_Analysis/ contains all code in R to process finalized datasets, ultimately generating Figure 2e, ExtData Fig 4f-g, and ExtData Fig 5a-h. OUTPUT/ contains final RDS files as output by each workflow, to use as input for each Analysis script. A reader may choose to skip the majority of the above workflow, and only download these files to immediately generate the Figures and/or further explore the data. In general, a reader can follow the A_import scripts to create initial Seurat objects, B_integrate scripts to integrate separate sequencing runs, C_annotate scripts to assign finalized cell types, and D_analysis scripts to generate all plots. Each workflow ostensibly only requires establishing certain library dependencies, as well as filepaths for INPUT and OUTPUT directories. Please note, several of the Integration scripts are particularly computationally intensive. These were successfully run by allocating 512 Gb of memory within our university’s supercomputing environment. As emphasized within code annotations, this entire workflow is run with R 4.3.2 and Seurat 5.1.0— our computation makes use of that FindClusters implementation, a shared nearest neighbor modularity optimization based clustering algorithm (Seurat 5.2 is not compatible). We specify the Leiden algorithm at these steps, which differs from the default Louvain algorithm, which offers certain computational advantages (see PMID 30914743). While results with Louvain would likely reach the same biological conclusions, the final clusters and annotations we provide here are only useful with Leiden implementation. Finally, we are grateful to our institution’s supercomputer admins, who helped establish a virtual environment with R-Reticulate, which allows for calling the Leiden algorithm (implemented in Python). Readers will similarly need to establish this dependency based on their local environment.

依赖Prdm16的抗原呈递细胞诱导肠道抗原免疫耐受 本仓库包含与上述论文相关的小鼠数据集的全部代码与计算分析内容。文件组织方式如下: 2024-04-26/ 目录包含Delta+7kb小鼠模型的原始测序数据。 2024-06-21/ 目录包含Rorc(t)_DeltaCD11c小鼠模型的原始测序数据。 2024-10-16/ 目录包含多组学(RNA+ATAC)原始测序数据。 2022AkagbosuEtAl/ 目录包含来自PubMed文献ID(PMID)36070798的公开多组学数据集,该数据集将与本研究的同类数据进行整合。(该数据集通过从基因表达综合数据库(Gene Expression Omnibus, GEO)仓库GSE174405下载,并使用CellRanger默认参数完成初始处理得到)。 2014-11-26/ 目录包含bigWig格式的批量RNA测序数据,作为Broad整合基因组浏览器(Integrative Genomics Viewer, IGV)工具的输入文件,最终用于生成图1b。 Delta7kB_Murine_Analysis/ 目录包含用于处理最终数据集的全部R语言代码,最终可生成图2a-c及扩展数据图4a-e。 Cd11c_Murine_Analysis/ 目录包含用于处理最终数据集的全部R语言代码,最终用于生成图2d。 Multiome_Murine_Analysis/ 目录包含用于处理最终数据集的全部R语言代码,最终可生成图2e、扩展数据图4f-g及扩展数据图5a-h。 OUTPUT/ 目录包含各分析流程生成的最终RDS文件,可作为各分析脚本的输入文件。研究者可跳过大部分前述流程,仅下载该目录下的文件即可直接生成图表或进一步探索数据集。 通常而言,研究者可按照A_import脚本创建初始Seurat对象,通过B_integrate脚本整合独立测序运行得到的数据,C_annotate脚本分配最终细胞类型,以及D_analysis脚本生成全部绘图结果。各流程仅需配置相应的库依赖项,并指定INPUT与OUTPUT目录的文件路径即可运行。 请注意,部分整合脚本的计算量极大。在本研究所属高校的超算环境中,通过分配512 Gb内存即可成功运行这些脚本。 如代码注释中强调的,本整套流程基于R 4.3.2与Seurat 5.1.0运行——本研究的计算使用了基于共享最近邻模块化优化的聚类算法FindClusters(Seurat 5.2版本不兼容)。在此步骤中我们指定使用Leiden算法,该算法相较于默认的Louvain算法具备一定的计算优势(详见PMID 30914743)。尽管使用Louvain算法也可得到一致的生物学结论,但本文提供的最终聚类与注释结果仅适配Leiden算法的实现。最后,我们感谢所在机构的超算管理员,他们协助搭建了包含R-Reticulate的虚拟环境,以支持调用Python实现的Leiden算法。研究者需根据本地环境配置相应的依赖项。

提供机构:
Zenodo
创建时间:
2025-04-14
二维码
社区交流群
二维码
科研交流群
商业服务