遇见数据集

Poplar_Isoform Expression_matrix_AND_Isoform_GTF

收藏
Figshare2020-04-07 更新2026-04-28 收录
官方服务:

资源简介:

A matrix of isoform expression values(FPKM) for each replicate samples for an unstructured population of 268 Populus deltoides, and a GTF file describing the transcriptome. Isoforms were discovered as follows:Three transcript assembly platforms were used in order to maximize isoform detection: (i) Cufflinks version 2.2 with parameters “--library-type fr-firststrand –u -F 0.05 --max-intron-length 12000 --no-faux-reads -g”; (ii) StringTie version 1.3.3 with parameters “-f 0.05 -j 2 –rf” , and (iii) Trinity version 2.3.2 in genome guided mode with parameters “--genome_guided_bam --genome_guided_max_intron 12000 --full_cleanup --SS_lib_type RF --min_contig_length 50”. The collection of Cufflinks and Stringtie isoforms detected for each sample were merged with Stringtie merge using parameters “-F 1 -f 0.05”. PASA version 2.0.2-r20151207 was used to reconcile this merged assembly and the assembly from Trinity using parameters “-C –R -t --cufflinks_gtf -I 12000 --ALT_SPLICE --ALIGNER gmap,blat”. Additionally, the assemblies generated by PASA were filtered by requiring that (i) all splice junctions be supported by at least 2 reads, and (ii) retained introns be supported by a median read coverage of at least 2 (Python scripts stored in github.com/jdLikesPlants/poplar_AS). Requiring a minimum read support for retained introns minimizes the possibility of incorrect identification of intron retention events from the sequencing of pre-mRNA. Finally, the filtered PASA assemblies for each sample were merged with Cuffmerge (Cufflinks version 2.2.1) to generate a master transcriptome that represents all of the potential AS events and transcript isoforms for the population. The resulting assembly was then reformatted and annotated using gffcompare version 0.9.9c (https://github.com/gpertea/gffread). This transcriptome was subjected to a secondary expression-based filtering pipeline to remove artifacts generated during the merge. Cufflinks version 2.2 was used in quantification mode (parameters: --library-type fr-firststrand -G -u -F 0.05 --max-intron-length 12000) to measure the expression of the transcripts in the merged assembly in each sample. To minimize the presence of incorrectly assembled transcripts in the merged transcriptome assembly, each transcript was required to be expressed above FPKM (fragments per kilobase of exon model per million reads mapped) 3 in at least two of three biological replicates of a given individual, and in at least 3 individuals in the population. This final merged and filtered transcriptome was used in all downstream analyses. It is included here as 'Polpar_deltoides_gffcmp.annotated_3geno_filt_cleaned.gtf. Additionally, any individual that did not have at least 15 million reads generated during sequencing in at least two of three replicates, as well as individuals for which only one replicate was sequenced were removed from analysis, resulting in a final set of 268 individuals.Quantification of expression of each isoform is provided in Poplar_Isoform_Expression_matrix.zip file.

本数据集包含268个个体组成的无群体结构的美洲杨(Populus deltoides)种群各生物学重复样本的异构体表达量矩阵(FPKM,每百万reads映射的外显子模型每千碱基片段数),以及描述该转录组的GTF文件。为最大化异构体检出效率,本研究采用了三种转录组组装平台:(i) Cufflinks v2.2,参数设置为"--library-type fr-firststrand –u -F 0.05 --max-intron-length 12000 --no-faux-reads -g";(ii) StringTie v1.3.3,参数设置为"-f 0.05 -j 2 –rf";(iii) 基因组引导模式下的Trinity v2.3.2,参数设置为"--genome_guided_bam --genome_guided_max_intron 12000 --full_cleanup --SS_lib_type RF --min_contig_length 50"。将每个样本检出的Cufflinks与StringTie异构体集合使用Stringtie merge工具进行合并,参数为"-F 1 -f 0.05"。使用PASA v2.0.2-r20151207,参数为"-C –R -t --cufflinks_gtf -I 12000 --ALT_SPLICE --ALIGNER gmap,blat",对上述合并后的组装结果与Trinity组装结果进行整合。对PASA生成的组装结果实施两轮过滤规则:① 所有剪接位点需至少有2条reads支持;② 保留内含子区域的中位测序覆盖度需至少为2(相关Python脚本存储于github.com/jdLikesPlants/poplar_AS)。为保留内含子事件设置最低read支持阈值,可有效降低因前体mRNA测序导致的内含子保留事件误判风险。将每个样本过滤后的PASA组装结果与Cuffmerge(Cufflinks v2.2.1)进行合并,生成覆盖该群体所有可变剪接(AS,Alternative Splicing)事件与转录异构体的核心转录组。使用gffcompare v0.9.9c(https://github.com/gpertea/gffread)对最终组装结果进行格式重构与注释。随后对该转录组实施基于表达量的二次过滤流程,以去除组装合并过程中产生的假阳性产物。采用Cufflinks v2.2的定量模式(参数:--library-type fr-firststrand -G -u -F 0.05 --max-intron-length 12000),对合并组装后的转录本在各样本中的表达量进行测定。为进一步剔除错误组装的转录本,设置如下筛选标准:每个转录本需在单个个体的3个生物学重复中至少2个重复的FPKM值高于3,且在该群体中至少3个个体中满足该表达量阈值。最终的合并过滤后转录组将用于所有下游分析,本数据集已包含该转录组文件:'Polpar_deltoides_gffcmp.annotated_3geno_filt_cleaned.gtf'。此外,对未在3个生物学重复中至少2个重复获得至少1500万条测序reads的个体,以及仅完成1个生物学重复测序的个体,均予以剔除,最终得到268个可用个体。各异构体的表达定量结果存储于Poplar_Isoform_Expression_matrix.zip压缩文件中。

创建时间:
2020-04-07
二维码
社区交流群
二维码
科研交流群
商业服务