Data for Indicine Cattle Multi-Assembly Pangenome Graph
收藏资源简介:
This dataset comprises a Minigraph–Cactus pangenome graph built from thirteen long-read de novo genome assemblies of five economically important Bos indicus breeds raised in Brazil (Nelore, Gyr, Guzerat, Sindi and Tabapuã), using the ARS-UCD2.0 taurine reference genome as backbone. The pangenome graph (IND_pangenome_final.gfa.gz) integrates the thirteen assemblies and captures approximately 120 Mb of non-reference sequence, including a 21.4 Mb core indicine genome shared across all five breeds, yielding a catalog of 114,141 structural variants. Individual genome assemblies (.fasta.gz) are also provided for each sampled animal, identified by breed and sample code, enabling independent reuse of the underlying genomic sequences beyond the pangenome graph. Citation Users of this dataset should cite the associated manuscript (Silva et al., Animal Genetics, in press) alongside this Zenodo record (DOI: 10.5281/zenodo.21535203). Funding This work was supported by the São Paulo Research Foundation (FAPESP), grants 2021/03101-9, 2023/06637-2, and 2024/18544-1, under the COWADAPT project: "Genetic variants for cattle adaptability to harsh environments uncovered through a bovine multi-assembly graph," a collaboration between the University of São Paulo (USP), ETH Zürich.



