Dataset 1: Genomic assemblies and annotations of Coffea species and subgenomes
收藏资源简介:
This dataset contains filtered genome assemblies and corresponding GFF3 annotation files for Coffea species. The original assemblies and annotations were obtained from public repositories: Coffea arabica ET-39 – NCBI: GCF_036785885.1Coffea arabica Caturra – NCBI: GCA_003713225.1Coffea arabica Bourbon – NCBI: GCA_030873655.1Coffea arabica Gesha – Zenodo: https://zenodo.org/records/10059814Coffea arabica Typica – Figshare: https://figshare.com/articles/dataset/b_A_chromosome-level_genome_assembly_of_b_b_Coffea_arabica_b_b_L_var_Kona_Typica_b/28425329/2Coffea eugenioides CCC68 – NCBI: GCA_003713205.1Coffea humblotiana – NCBI: GCA_023065735.1Coffea canephora – NCBI: GCA_900059795.1 After obtaining the assemblies from public repositories, filters were applied to retain only chromosomes and contigs ≥500 kb for all species and cultivars. For Coffea arabica (cultivars Gesha, Caturra, Bourbon, ET-39, and Typica), the assemblies were additionally separated into subgenomes. The corresponding GFF3 annotation files were also filtered to match the retained chromosomes/contigs. Subsequently, the filtered GFF3 files were processed with the AGAT toolkit to generate new longest isoform GFF3 annotation files. Using these longest GFF3 files together with the filtered genomes, protein and CDS sequences were extracted in AGAT. For Coffea arabica Bourbon and Coffea humblotiana, no GFF3 annotations were available, but their genome assemblies were filtered in the same way as for the other species. Genome file extension = .fasta | Annotation file extension = .gff3 | Protein file extension = .faa | CDS file extension = .fna Scripts used for AGAT processing are available at: https://github.com/daisysotero/Coffea-analyses-2025 sgC = canephora-derived subgenome | sgE = eugenioides-derived subgenome



