遇见数据集

AmarylOmicBase: An integrated transcriptome dataset for comparative analysis of Amaryllidoideae species

收藏
Zenodo2026-05-28 更新2026-05-29 收录
官方服务:

资源简介:

General description Compiled dataset of transcriptome assemblies, transcriptome annotation and expression quantification for 27 Amaryllidoideae species and 4 hybrid cultivars: Amaryllis belladonna, Clivia miniata, Crinum asiaticum, Crinum x powellii, Galanthus elwesii, Galanthus sp., Hippeastrum cv. Blossom Peacock, Hippeastrum cv. Jewel, Hippeastrum cv. Royal Velvet, Hippeastrum striatum, Hippeastrum vittatum, Leucojum aestivum, Lycoris aurea, Lycoris chinensis, Lycoris incarnata, Lycoris longituba, Lycoris radiata, Lycoris sprengeri, Narcissus aff pseudonarcissus, Narcissus papyraceus, Narcissus pseudonarcissus, Narcissus cv. Tête-à-Tête, Narcissus tazetta, Narcissus viridiflorus, Phycella aff cyrtanthoides, Rhodophiala pratensis, Scadoxus multiflorus, Traubia modesta, Zephyranthes candida, Zephyranthes carinata, Zephyranthes treatiae. Data includes transcriptome assemblies, Transdecoder predictions of peptide sequences (along with predicted coding seqeunces, GFF3 and BED files), expression quantification (count and TPM matrices for genes and trinity isoforms obtained with Kallisto and from Salmon), and annotation results (from EggNOG, Pfam, Uniprot Swissprot and Rfam), as well as signal peptide and transmemberane domain predictions (from SignalP and TmHMM). Annotations were compiled into a report for each species using Trinotate. For Lycoris aurea and Narcissus cv. Tête-à-Tête, there are two assemblies and corresponding files: Lycoris_aurea_PB and Lycoris_aurea_TH, and Narcissus_TaT_PB Narcissus_TaT_TH. PB assemblies were constructed solely with long-read sequencing data (PacBio, PB); while TH assemblies were constructed with short reads using Trinity, with long-read assembly being used for scaffolding step of Trinity (Trinity Hybrid, TH). Expression quantification was published on NCBI GEO (accessions GSE329951, GSE329957, GSE330014 and GSE331457). File descriptions All unitigs and predicted protein sequences are prefixed with an acronym to identify the species: Species Acronym NCBI TSA accession Amaryllis belladonna Ambel deposited on GSE331457 Clivia miniata Clmin DBNKRK000000000 Crinum asiaticum Crasi DBNIJL000000000 Crinum x powellii Crpow DBNKRS000000000 Galanthus elwesii Gaelw DBNKRO000000000 Galanthus sp. Gasp DBNKRR000000000 Hippeastrum cv. Blossom Peacock HispBP DBNMYE000000000 Hippeastrum cv. Jewel HispJW DBNMYC000000000 Hippeastrum cv. Royal Velvet HispRV DBNMYD000000000 Hippeastrum striatum Histr DBNIJG000000000 Hippeastrum vittatum Hivit DBNKRF000000000 Leucojum aestivum Leaes DBNIJJ000000000 Lycoris aurea (PB) Lyaur DBNFTY000000000 Lycoris aurea (TH) Lyaur DBNKRI000000000 Lycoris chinensis Lychi DBNKRE000000000 Lycoris incarnata Lyinc DBNKRD000000000 Lycoris longituba Lylon DBNKRG000000000 Lycoris radiata Lyrad deposited on GSE331457 Lycoris sprengeri Lyspr DBNKRL000000000 Narcissus aff pseudonarcissus Naafps DBNKRP000000000 Narcissus papyraceus Napap DBNKRN000000000 Narcissus pseudonarcissus Napse DBNIJK000000000 Narcissus Tête-à-Tête (PB) NaspPB DBNNFO000000000 Narcissus Tête-à-Tête (TH) NaspTH DBNUFN000000000 Narcissus tazetta Nataz DBNPME000000000 Narcissus viridiflorus Navir deposited on GSE331457 Phycella sp. Phsp deposited on GSE331457 Rhodophiala pratensis Rhpra deposited on GSE331457 Scadoxus multiflorus Scmul DBNIJH000000000 Traubia modesta Trmod deposited on GSE331457 Zephyranthes candida Zecan DBNKRH000000000 Zephyranthes carinata Zecar DBNIJI000000000 Zephyranthes treatiae Zetre deposited on GSE331457 Files Files are organized by type of analysis/data, meaning all expression quantification data obtained with Kallisto are compressed into a single file, all results from BlastP are in the same file, etc. Compressed file name Individual file type Content Blastp_Uniprot.tar.gz Tabular 33 files (1 per assembly) generated with BLASTP against Uniprot SwissProt release 2024_04. Files in blast output format 6 (standard columns). EggNOG_Emapper.tar.gz Tabular 33 tabular files (1 per assembly). Generated with eggnog.emapper (annotations file format described in eggnog.emapper’s wiki) Expression_Kallisto.tar.gz Tabular 128 files (4 per assembly, 2 for long-read assemblies): read counts and TPM values for both trinity “genes” and trinity “isoforms”). Row names are “gene”/ “isoform” IDs, column names are SRA run IDs. Expression_Salmon.tar.gz Tabular 128 files (4 per assembly, 2 for long-read assemblies): read counts and TPM values for both trinity “genes” and trinity “isoforms”). Row names are “gene”/ “isoform” IDs, column names are SRA run IDs. Final_assemblies.tar.gz Fasta 33 files (1 per assembly). Assemblies generated in this study, after contamination and expression filtering. HMMScan_PfamA.tar.gz Domain hits table 33 files (1 per assembly). Generated with hmmscan (option ‑‑domtblout, domain hits table, explained in hmmer’s user guide [119]) against the Pfam-A database. Infernal_Rfam.tar.gz Target hits table format 2 33 files (1 per assembly) generated with Infernal's cmscan (using the Trinotate wrapper) against the Rfam database. Table format 2 described in Infernal's user guide section 6). Metadata_studies.tar.gz Tabular (semi-colon separated columns) 31 tabular files (one per species/cultivar) indicating: Bioproject ID, SRA run ID, Sample name, Biosample ID, tissue, genotype (cultivar, when specified), treatment, batch, original publication citation, and DOI of original publication. Signalp6.tar.gz Tabular or GFF3 99 files (3 per assembly: prediction_results.txt, output.gff3 and region_output.gff3) generated with SignalP6. TmHMM2.tar.gz Tabular 33 files (1 per assembly) generated with TmHMM2 (short format, described in the guide tab of https://services.healthtech.dtu.dk/services/TMHMM-2.0/ Transcriptomes_unfiltered.tar.gz Fasta 33 files (1 per assembly). Unfiltered (prior to contamination and expression screening) assemblies generated in this study. Transdecoder_bed.tar.gz BED 33 BED files (1 per assembly) generated with TransDecoder.Predict (using BLASTp against Uniprot and hmmscan against Pfam) Transdecoder_cds.tar.gz Fasta 33 fasta files (1 per assembly) of predicted coding sequences generated with TransDecoder.Predict (using BLASTp against Uniprot and hmmscan against Pfam) Transdecoder_gff3.tar.gz GFF3 33 GFF3 files (1 per assembly) generated with TransDecoder.Predict (using BLASTp against Uniprot and hmmscan against Pfam). Transdecoder_proteomes.tar.gz Fasta 33 predicted proteome files (.pep, 1 per assembly) generated with TransDecoder.Predict (using BLASTp against Uniprot and hmmscan against Pfam).

提供机构:
Zenodo
创建时间:
2026-05-27
二维码
社区交流群
二维码
科研交流群
商业服务