Data for Long-read RNA sequencing identifies region- and sex-specific C57BL/6J mouse brain mRNA isoform expression and usage
收藏资源简介:
data_minus_bam.tar.gz contains all files from the data directory (except for bam outputs) associated with the 230227_EJ_MouseBrainIsoDiv GitHub project and includes the following: - comparison_gene_lists/: The RData in the following directory contains all comparison gene lists with DGE, DTE, and DTU for importing into the R environment and reproducing analyses. - all_comparison_gene_lists.Rdata - cpm_out/: The RData in the following directory contains the processed counts per million and formatted metadata for downstream analyses. - cpm_counts_metadata.RData - deseq2_data/: All files in the following directory are Rds files with deseq2 results for the study design indicated in the file name. If the file name includes “gene” it was done at the gene level and “transcript” indicates the analysis was done at the transcript level. If a filename includes two regions, it is a comparison between the two, a file name with one region denotes either “one vs all” or “male vs female”. Any filename that includes “sex” is male vs female in the indicated region(s). - all_regions_sex_gene_results.Rds - all_regions_sex_transcript_results.Rds - cerebellum_cortex_results.Rds - cerebellum_cortex_transcripts_results.Rds - cerebellum_gene_results.Rds - cerebellum_hippocampus_results.Rds - cerebellum_hippocampus_transcripts_results.Rds - cerebellum_sex_gene_results.Rds - cerebellum_sex_transcript_results.Rds - cerebellum_striatum_results.Rds - cerebellum_striatum_transcripts_results.Rds - cerebellum_transcript_results.Rds - cortex_gene_results.Rds - cortex_hippocampus_results.Rds - cortex_hippocampus_transcripts_results.Rds - cortex_sex_gene_results.Rds - cortex_sex_transcript_results.Rds - cortex_striatum_results.Rds - cortex_striatum_transcripts_results.Rds - cortex_transcript_results.Rds - hippocampus_gene_results.Rds - hippocampus_sex_gene_results.Rds - hippocampus_sex_transcript_results.Rds - hippocampus_striatum_transcripts_results.Rds - hippocampus_transcript_results.Rds - striatum_gene_results.Rds - striatum_hippocampus_results.Rds - striatum_hippocampus_transcripts_results.Rds - striatum_sex_gene_results.Rds - striatum_sex_transcript_results.Rds - striatum_transcript_results.Rds - gencode_annotations/: This directory contains the exact GENCODE genome and transcriptome annotations used for our analyses - GRCm39.primary_assembly.genome.fa - GRCm39.primary_assembly.genome.fa.fai - gencode.vM31.primary_assembly.annotation.gtf - gffread/: This directory contains the generated fasta files with exact isoform sequences for novel and annotated genes required for creating isoformSwitchAnalyzeR objects. - isoform_sequences.fa - isoform_sequences_linear.fa - nextflow/: All files in the following directories in the overarching nextflow are direct outputs from the nf-core nanoseq pipeline. For specific information on nanoseq pipeline outputs, please refer to https://nf-co.re/nanoseq/3.1.0/docs/output - bambu/ - counts_gene.txt - counts_transcript.txt - extended_annotations.gtf - extended_annotations.gtf.idx - versions.yml - fastqc/ - There are 2 files for each of the 40 samples. Below is a representative example, of files expected for each of the 40 samples: - sample01_R1_1_fastqc.html - sample01_R1_1_fastqc.zip - minimap2/ - bam/ This directory has been removed to save space, please contact us for more information. - bigBed/ - There is 1 file for each of the 40 samples. Below is a representative example, of files expected for each of the 40 samples: - sample01_R1.bigBed - bigWig/ - There is 1 file for each of the 40 samples. Below is a representative example, of files expected for each of the 40 samples: - sample01_R1.bigWig - genome/ - GRCm39.primary_assembly.genome.fa.mmi - samtools_stats/ - There are 3 files for each of the 40 samples. Below is a representative example, of files expected for each of the 40 samples: - sample01_R1.sorted.bam.flagstat - sample01_R1.sorted.bam.idxstats - sample01_R1.sorted.bam.stats - multiqc/ - multiqc_data/ - mqc_samtools-idxstats-mapped-reads-plot_Normalised_Counts.txt - mqc_samtools-idxstats-mapped-reads-plot_Observed_over_Expected_Counts.txt - mqc_samtools-idxstats-mapped-reads-plot_Raw_Counts.txt - mqc_samtools-idxstats-xy-plot_1.txt - mqc_samtools_alignment_plot_1.txt - multiqc.log - multiqc_data.json - multiqc_general_stats.txt - multiqc_samtools_flagstat.txt - multiqc_samtools_idxstats.txt - multiqc_samtools_stats.txt - multiqc_sources.txt - multiqc_plots/ - pdf/ - mqc_samtools-idxstats-mapped-reads-plot_Normalised_Counts.pdf - mqc_samtools-idxstats-mapped-reads-plot_Observed_over_Expected_Counts.pdf - mqc_samtools-idxstats-mapped-reads-plot_Raw_Counts.pdf - mqc_samtools-idxstats-xy-plot_1.pdf - mqc_samtools-idxstats-xy-plot_1_pc.pdf - mqc_samtools_alignment_plot_1.pdf - mqc_samtools_alignment_plot_1_pc.pdf - png/ - *The same multiqc plots as the pdf directory, but in png format* - svg/ - *The same multiqc plots as the pdf and png directory, but in svg format* - multiqc_report.html - versions.yml - nanoplot/ - fastq/ - Contains 40 directories for 40 samples, each containing 12 files obtained from running nanoplot with the nf-core nanoseq pipeline. Below is a representative example, but this repo contains 1 directory per sample: - sample01_R1/ - Dynamic_Histogram_Read_length.html - HistogramReadlength.png - LengthvsQualityScatterPlot_dot.png - LengthvsQualityScatterPlot_kde.png - LogTransformed_HistogramReadlength.png - NanoPlot-report.html - NanoPlot_20230413_1600.log - NanoPlot_20230413_2047.log - NanoStats.txt - Weighted_HistogramReadlength.png - Weighted_LogTransformed_HistogramReadlength.png - Yield_By_Length.png - pipeline_info/ - execution_report_2023-04-13_15-46-11.html - execution_timeline_2023-04-13_15-46-11.html - execution_trace_2023-04-13_10-59-24.txt - execution_trace_2023-04-13_15-46-11.txt - pipeline_dag_2023-04-13_15-46-11.svg - samplesheet.valid.csv - software_versions.yml - switchlist_fasta/: This directory contains the generated fasta files for amino acids and nucleotides for individual isoformSwitchAnalyzeR objects required for downstream analyses. - cerebellum_AA.fasta - cerebellum_nt.fasta - cerebellum_sex_AA.fasta - cerebellum_sex_nt.fasta - cortex_AA.fasta - cortex_nt.fasta - cortex_sex_AA.fasta - cortex_sex_nt.fasta - hippocampus_AA.fasta - hippocampus_nt.fasta - region_region_AA.fasta - region_region_nt.fasta - striatum_AA.fasta - striatum_nt.fasta - striatum_sex_AA.fasta - striatum_sex_nt.fasta - switchlist_objects/: This directory contains intermediate and final isoformSwitchAnalyzeR objects. “Region_all” in the filename is a list of four switchlists that compare a single brain region (cerebellum, cortex, hippocampus, striatum) to all others in aggregate. “Region_sex” in the filename is a list of four switchlists (cerebellum, cortex, hippocampus, striatum) that compare across sexes (male and female). “Region_region” denotes a single switchlist that includes all pairwise region comparisons. “Sex” in the name without “region” is comparing all regions in aggregate. - de_added/: This directory contains final isoformSwitchAnalyzeR objects that include open reading frame and differential expression results incorporated. - region_all_switchlist_list_orf_de.Rds - region_region_orf_de.Rds - region_sex_switchlist_list_orf_de.Rds - orf_added/: This directory contains intermediate and final isoformSwitchAnalyzeR objects with open reading frame information added. - region_all_switchlist_list.Rds - region_region_switchlist_analyzed.Rds - region_sex_switchlist_list.Rds - sex_switchlist_analyzed.Rds - pfam_added/: This directory contains final isoformSwitchAnalyzeR objects (including de and orf information) with added protein domain information. Please note pfam does not comprehensively identify all protein domains for every gene. - region_all_list_orf_de_pfam.Rds - region_region_orf_de_pfam.Rds - region_sex_list_orf_de_pfam.Rds - raw/: This directory contains the initial isoformSwitchAnalyzeR objects, without additional information added. - region_all_switchlist_list.Rds - region_region_switchlist_analyzed.Rds - region_sex_switchlist_list.Rds - sex_switchlist.Rds
data_minus_bam.tar.gz 包含了与230227_EJ_MouseBrainIsoDiv GitHub项目相关的data目录下全部文件(BAM输出结果除外),具体包含以下内容: - comparison_gene_lists/:该目录下的RData文件包含所有带差异基因表达(Differential Gene Expression, DGE)、差异转录本表达(Differential Transcript Expression, DTE)以及差异转录本使用(Differential Transcript Usage, DTU)的比对基因列表,可导入R环境并复现分析流程。 - all_comparison_gene_lists.Rdata - cpm_out/:该目录下的RData文件包含处理后的每百万读取数(CPM)与格式化后的元数据,用于下游分析。 - cpm_counts_metadata.RData - deseq2_data/:该目录下所有文件均为Rds格式文件,包含对应文件名中标注的实验设计的DESeq2分析结果。若文件名包含“gene”,则代表该分析基于基因水平进行;若包含“transcript”,则代表分析基于转录本水平进行。若文件名包含两个区域,则代表该文件为两个区域间的比对结果;若仅包含一个区域,则代表分析为“单区域vs其余所有区域”或“雄性vs雌性”。文件名中包含“sex”的文件,均为指定区域内的雄性与雌性比对结果。 - all_regions_sex_gene_results.Rds - all_regions_sex_transcript_results.Rds - cerebellum_cortex_results.Rds - cerebellum_cortex_transcripts_results.Rds - cerebellum_gene_results.Rds - cerebellum_hippocampus_results.Rds - cerebellum_hippocampus_transcripts_results.Rds - cerebellum_sex_gene_results.Rds - cerebellum_sex_transcript_results.Rds - cerebellum_striatum_results.Rds - cerebellum_striatum_transcripts_results.Rds - cerebellum_transcript_results.Rds - cortex_gene_results.Rds - cortex_hippocampus_results.Rds - cortex_hippocampus_transcripts_results.Rds - cortex_sex_gene_results.Rds - cortex_sex_transcript_results.Rds - cortex_striatum_results.Rds - cortex_striatum_transcripts_results.Rds - cortex_transcript_results.Rds - hippocampus_gene_results.Rds - hippocampus_sex_gene_results.Rds - hippocampus_sex_transcript_results.Rds - hippocampus_striatum_transcripts_results.Rds - hippocampus_transcript_results.Rds - striatum_gene_results.Rds - striatum_hippocampus_results.Rds - striatum_hippocampus_transcripts_results.Rds - striatum_sex_gene_results.Rds - striatum_sex_transcript_results.Rds - striatum_transcript_results.Rds - gencode_annotations/:该目录包含本次分析所用的官方GENCODE基因组与转录组注释文件。 - GRCm39.primary_assembly.genome.fa - GRCm39.primary_assembly.genome.fa.fai - gencode.vM31.primary_assembly.annotation.gtf - gffread/:该目录包含生成的FASTA格式文件,内含用于构建isoformSwitchAnalyzeR对象所需的新基因与注释基因的精确转录本序列。 - isoform_sequences.fa - isoform_sequences_linear.fa - nextflow/:该总目录下所有子目录的文件均为nf-core nanoseq流程的直接输出结果。如需了解nanoseq流程输出的详细信息,请访问:https://nf-co.re/nanoseq/3.1.0/docs/output - bambu/: - counts_gene.txt - counts_transcript.txt - extended_annotations.gtf - extended_annotations.gtf.idx - versions.yml - fastqc/:40个样本各对应2个文件,以下为单个样本的典型文件示例: - sample01_R1_1_fastqc.html - sample01_R1_1_fastqc.zip - minimap2/: - bam/:该目录已被移除以节省存储空间,如需获取更多信息请联系我们。 - bigBed/:40个样本各对应1个文件,以下为单个样本的典型文件示例: - sample01_R1.bigBed - bigWig/:40个样本各对应1个文件,以下为单个样本的典型文件示例: - sample01_R1.bigWig - genome/: - GRCm39.primary_assembly.genome.fa.mmi - samtools_stats/:40个样本各对应3个文件,以下为单个样本的典型文件示例: - sample01_R1.sorted.bam.flagstat - sample01_R1.sorted.bam.idxstats - sample01_R1.sorted.bam.stats - multiqc/: - multiqc_data/: - mqc_samtools-idxstats-mapped-reads-plot_Normalised_Counts.txt - mqc_samtools-idxstats-mapped-reads-plot_Observed_over_Expected_Counts.txt - mqc_samtools-idxstats-mapped-reads-plot_Raw_Counts.txt - mqc_samtools-idxstats-xy-plot_1.txt - mqc_samtools_alignment_plot_1.txt - multiqc.log - multiqc_data.json - multiqc_general_stats.txt - multiqc_samtools_flagstat.txt - multiqc_samtools_idxstats.txt - multiqc_samtools_stats.txt - multiqc_sources.txt - multiqc_plots/: - pdf/: - mqc_samtools-idxstats-mapped-reads-plot_Normalised_Counts.pdf - mqc_samtools-idxstats-mapped-reads-plot_Observed_over_Expected_Counts.pdf - mqc_samtools-idxstats-mapped-reads-plot_Raw_Counts.pdf - mqc_samtools-idxstats-xy-plot_1.pdf - mqc_samtools-idxstats-xy-plot_1_pc.pdf - mqc_samtools_alignment_plot_1.pdf - mqc_samtools_alignment_plot_1_pc.pdf - png/:与pdf目录存储完全一致的质控绘图结果,格式为PNG - svg/:与pdf、png目录存储完全一致的质控绘图结果,格式为SVG - multiqc_report.html - versions.yml - nanoplot/: - fastq/:包含40个样本对应的40个目录,每个目录存储通过nf-core nanoseq流程运行nanoplot生成的12个文件,以下为单个样本目录的典型示例: - sample01_R1/: - Dynamic_Histogram_Read_length.html - HistogramReadlength.png - LengthvsQualityScatterPlot_dot.png - LengthvsQualityScatterPlot_kde.png - LogTransformed_HistogramReadlength.png - NanoPlot-report.html - NanoPlot_20230413_1600.log - NanoPlot_20230413_2047.log - NanoStats.txt - Weighted_HistogramReadlength.png - Weighted_LogTransformed_HistogramReadlength.png - Yield_By_Length.png - pipeline_info/: - execution_report_2023-04-13_15-46-11.html - execution_timeline_2023-04-13_15-46-11.html - execution_trace_2023-04-13_10-59-24.txt - execution_trace_2023-04-13_15-46-11.txt - pipeline_dag_2023-04-13_15-46-11.svg - samplesheet.valid.csv - software_versions.yml - switchlist_fasta/:该目录包含生成的氨基酸与核苷酸FASTA格式文件,用于构建下游分析所需的单个isoformSwitchAnalyzeR对象。 - cerebellum_AA.fasta - cerebellum_nt.fasta - cerebellum_sex_AA.fasta - cerebellum_sex_nt.fasta - cortex_AA.fasta - cortex_nt.fasta - cortex_sex_AA.fasta - cortex_sex_nt.fasta - hippocampus_AA.fasta - hippocampus_nt.fasta - region_region_AA.fasta - region_region_nt.fasta - striatum_AA.fasta - striatum_nt.fasta - striatum_sex_AA.fasta - striatum_sex_nt.fasta - switchlist_objects/:该目录包含中间产物与最终的isoformSwitchAnalyzeR对象。文件名中含“Region_all”的文件,为将单个脑区(小脑、皮层、海马体、纹状体)与其余所有脑区汇总后的4个切换列表的集合;文件名含“Region_sex”的文件,为包含4个脑区(小脑、皮层、海马体、纹状体)的跨性别(雄性vs雌性)切换列表的集合;“Region_region”代表包含所有脑区两两比对结果的单个切换列表;文件名不含“region”的“Sex”相关文件,为所有脑区汇总后的跨性别比对结果。 - de_added/:包含整合了开放阅读框(Open Reading Frame, ORF)与差异表达(Differential Expression, DE)结果的最终isoformSwitchAnalyzeR对象。 - region_all_switchlist_list_orf_de.Rds - region_region_orf_de.Rds - region_sex_switchlist_list_orf_de.Rds - orf_added/:包含添加了开放阅读框(Open Reading Frame, ORF)信息的中间与最终isoformSwitchAnalyzeR对象。 - region_all_switchlist_list.Rds - region_region_switchlist_analyzed.Rds - region_sex_switchlist_list.Rds - sex_switchlist_analyzed.Rds - pfam_added/:包含添加了蛋白质结构域信息的最终isoformSwitchAnalyzeR对象(已整合差异表达与开放阅读框信息)。需注意,Pfam无法全面识别所有基因的全部蛋白质结构域。 - region_all_list_orf_de_pfam.Rds - region_region_orf_de_pfam.Rds - region_sex_list_orf_de_pfam.Rds - raw/:包含初始的isoformSwitchAnalyzeR对象,未添加任何额外信息。 - region_all_switchlist_list.Rds - region_region_switchlist_analyzed.Rds - region_sex_switchlist_list.Rds - sex_switchlist.Rds



