De novo transcriptome and annotation files for Strombocarpa tamarugo
收藏资源简介:
This transcriptome was assembled from 4 Illumina NovaSeq 6000 libraries generated as 2 × 150 bp paired-end reads. These libraries were produced from RNA samples extracted from leaves of Strombocarpa tamarugo in the field (further details available at Gajardo et al. (2024) - https://doi.org/10.1007/s00425-024-04484-1) Raw reads were quality-filtered with fastp v1.0.1 to remove low-quality bases, excessive Ns, adapter contamination, and poly(A) tails. Filtered reads were then assembled independently using Trinity v2.15.2, SOAPdenovo v1.0.4, Trans-ABySS v2.0.1, RNA-Bloom v2.0.1, and rnaSPAdes v4.2.0, all with default parameters. The resulting assemblies were concatenated and subjected to successive redundancy-reduction steps, including exact sequence deduplication, transcript length filtering (minimum 300 bp), high-stringency clustering with CD-HIT v4.8.1 at 99% sequence identity, containment removal, and expression-based filtering using Salmon v1.10.3. Only transcripts meeting the minimum expression threshold were retained. Candidate open reading frames (ORFs) were identified with TransDecoder 2 (github.com/Markusjsommer/TD2; Mao et al., 2025 - https://www.biorxiv.org/content/10.1101/2025.04.13.648579v1), and the longest predicted peptides were searched against a custom Viridiplantae-filtered UniRef100 database using MMseqs2 (commit befcb1130a73ddba247fe3eaa504e7ffb790d150). This database was generated by downloading UniRef100 and its associated taxonomy information through MMseqs2 and filtering entries assigned to Viridiplantae. These homology results were used as supporting evidence to retain high-confidence ORFs in TransDecoder 2. To further reduce transcriptome redundancy, transcripts were searched against the proteome of closely related species Neltuma alba using DIAMOND blastx v2.0.15, and the top two matching ORFs were retained. ORFs without matches to the sister-species proteome were retained as putatively novel only if they showed strong expression support (≥100 total mapped reads across all samples) and encoded peptides of at least 150 amino acids. The transcripts and predicted proteins were further annotated with hmmscan (HMMER v3.4) against the Pfam-A database (release 13-01-2026), and the Mercator v4.8 database (Schwacke et al., 2019 - https://doi.org/10.1016/j.molp.2019.01.003). Additional fasta and text manipulation steps, including duplicate removal and identifier standardization, were performed using SeqKit v2.3.0, Seqtk v1.4-r122, BBMap v39.33, and standard Bash command-line utilities. The script with commands to generate the files in this dataset are in 10.5281/zenodo.18874676.
本转录组由4个Illumina NovaSeq 6000测序文库组装得到,文库采用2×150 bp双端测序读段构建。该类文库的RNA样本提取自野外生境下的Strombocarpa tamarugo叶片,详细信息可参见Gajardo等(2024)的研究(DOI:https://doi.org/10.1007/s00425-024-04484-1)。 原始测序读段使用fastp v1.0.1进行质量过滤,以去除低质量碱基、过量的模糊碱基(N)、接头污染及poly(A)尾序列。随后,使用Trinity v2.15.2、SOAPdenovo v1.0.4、Trans-ABySS v2.0.1、RNA-Bloom v2.0.1及rnaSPAdes v4.2.0这5款工具以默认参数分别对过滤后的读段进行独立组装。 将各工具得到的组装结果进行拼接后,依次执行一系列冗余去除步骤:包括精确序列去重、转录本长度过滤(最小长度阈值为300 bp)、使用CD-HIT v4.8.1以99%序列相似度进行高严格性聚类、移除序列包含型冗余,以及通过Salmon v1.10.3进行表达量过滤。仅保留满足最低表达量阈值的转录本。 使用TransDecoder 2(源码地址:github.com/Markusjsommer/TD2;Mao等,2025,DOI:https://www.biorxiv.org/content/10.1101/2025.04.13.648579v1)预测候选开放阅读框(open reading frames, ORFs);随后使用MMseqs2(提交版本:befcb1130a73ddba247fe3eaa504e7ffb790d150),将最长的预测肽段与自定义的Viridiplantae过滤版UniRef100数据库进行比对。该自定义数据库通过MMseqs2下载UniRef100及其关联分类学信息,并筛选出分类为Viridiplantae的条目构建得到。上述同源比对结果作为支持证据,用于保留TransDecoder 2预测的高可信度ORFs。 为进一步降低转录组冗余度,使用DIAMOND blastx v2.0.15将转录本与近缘物种Neltuma alba的蛋白质组进行比对,仅保留比对得分前两位的ORFs。对于未在该近缘物种蛋白质组中找到比对结果的ORFs,仅当其满足以下两个条件时方可作为推定新序列保留:一是具备较强的表达支持(所有样本中总比对读段数≥100),二是编码的肽段长度至少为150个氨基酸。 使用hmmscan(HMMER v3.4)分别针对Pfam-A数据库(发布版本:2026年1月13日)及Mercator v4.8数据库(Schwacke等,2019,DOI:https://doi.org/10.1016/j.molp.2019.01.003),对转录本及预测蛋白质序列进行功能注释。 其余FASTA格式文件及文本处理步骤(包括重复序列移除、标识符标准化等)通过SeqKit v2.3.0、Seqtk v1.4-r122、BBMap v39.33及标准Bash命令行工具完成。 用于生成本数据集所有文件的命令脚本可通过DOI:10.5281/zenodo.18874676获取。



