Dataset 2: Genomic assemblies and annotations of Coffea species and subgenomes
收藏资源简介:
This dataset contains genome assemblies and corresponding annotation files for various Coffea species. The genomic data were retrieved from public repositories and underwent standardization and validation to ensure consistency across files. Genome assemblies were filtered to retain chromosomes and contigs ≥500 kb. For Coffea arabica, assemblies were further separated into subgenomes derived from Coffea canephora and Coffea eugenioides. Filtered GFF3 files matching these sequences were processed with the AGAT toolkit to generate longest isoform annotations, which were then used together with the filtered genomes to extract protein and CDS sequences. Scripts used for AGAT processing are available at: https://github.com/daisysotero/Coffea-analyses-2025 *Genome file extension = *_500kb.fasta | Annotation file extension = *_500kb.gff3 | Protein file extension = *_500kb_prot.fasta | sgCC or sgC = canephora-derived subgenome | sgEE or sgE = eugenioides-derived subgenome



