遇见数据集

Dataset 2: Genomic assemblies and annotations of Coffea species and subgenomes

收藏
Zenodo2026-02-06 更新2026-05-26 收录
官方服务:

资源简介:

This dataset contains genome assemblies and corresponding annotation files for various Coffea species. The genomic data were retrieved from public repositories and underwent standardization and validation to ensure consistency across files. Genome assemblies were filtered to retain chromosomes and contigs ≥500 kb. For Coffea arabica, assemblies were further separated into subgenomes derived from Coffea canephora and Coffea eugenioides. Filtered GFF3 files matching these sequences were processed with the AGAT toolkit to generate longest isoform annotations, which were then used together with the filtered genomes to extract protein and CDS sequences. Scripts used for AGAT processing are available at: https://github.com/daisysotero/Coffea-analyses-2025 *Genome file extension = *_500kb.fasta | Annotation file extension = *_500kb.gff3 | Protein file extension = *_500kb_prot.fasta | sgCC or sgC = canephora-derived subgenome | sgEE or sgE = eugenioides-derived subgenome

提供机构:
Zenodo
创建时间:
2026-02-06
二维码
社区交流群
二维码
科研交流群
商业服务