遇见数据集

Genome feature files (GFF3) for the eight 'NIAB Elite MAGIC' wheat founders

收藏
Zenodo2025-12-11 更新2026-05-26 收录
官方服务:

资源简介:

Genome Assembly Initially, scaffold level assemblies for the eight 'NIAB Elite MAGIC' wheat (Triticum aestivum L.) founders (Mackay et al., 2014) were generated as described in the w2rap pipeline (https://github.com/bioinfologics/w2rap). First, the w2rap-contigger (https://github.com/bioinfologics/w2rap-contigger, branch: bj) was used to generate contigs from paired-end reads using k=200 and all other parameters left as default. To scaffold the contigs, mate pair libraries were produced for each line. Mate pairs were processed and filtered as described in the w2rap pipeline documentation. The w2rap version of the SOAP scaffolder was used to perform scaffolding with K=71 for s_prepare and k=71 for s_map with the other parameters left as default. The SOAP config file used max_rd_len=900 with paired-end reads used as the first library to scaffold, followed by the mate-pair libraries in increasing insert size order. The number of links to join sequences was left at the default of 3 for paired-end reads and 5 for mate-pair reads. Scaffolds shorter than 500 bp were removed from the final assemblies. Pseudomolecule construction: Reference-guided pseudomolecules were constructed with a modified TRITEX pipeline (Monat et al., 2019). A guide map was derived from the chromosome-scale sequence of the 'winter' German bread wheat cv. Julius (Walkowiak et al., 2020). Single-copy regions were extracted from the Julius assembly using BBDuk (Bushnell et al., 2014) as described by Jayakodi et al. (2020) and aligned to the W2RAP assemblies using Minimap2 (Li, 2018). Alignments with mapping quality below 30 and alignment lengths less than 2 kb were filtered. W2RAP contigs longer than 300 kb with at least 10 kb aligned single-copy sequences were assigned to chromosomal locations and oriented following majority rules. AGP files specifying the order and orientation of contigs were written and compiled into FASTA sequences using TRITEX functions (https://tritexassembly.bitbucket.io). Hi-C was performed according to the in-situ Hi-C protocol of Padmarasu et al. (2019). Hi-C reads were aligned to the W2RAP contigs using the TRITEX pipeline using Minimap2 for alignment, Novosort for sorting (http://www.novocraft.com/products/novosort/) and SAMtools (Danecek et al., 2021) and BEDTools (Quinlan & Hall, 2010) for aggregation of information. Hi-C contact maps at 1 Mb resolution arranged according to the chromosomal AGP files were plotted with TRITEX functions and manually inspected for off-diagonal signals to spot large structural variants between the assembled genomes and the Julius guide genome. Annotation Genome annotations were performed on CropDiversity-HPC, described by Percival-Alwyn et al. (2024). Gene-level annotations were generated by lifting over the IWGSC RefSeq v1.1 Chinese Spring wheat reference annotation (Alaux et al., 2018; IWGSC, 2018) using Liftoff (Shumate & Salzberg, 2021). The following Liftoff options were specified: -flank 0.05, -exclude_partial, -copies, -polish, -chroms, and -unplaced. Minimap2 (Li, 2018) was used as the aligner with the options: -a --end-bonus 5 --eqx -N 50 -p 0.5 -I 20G. Resulting GFF files were filtered to remove duplicate gene annotations (i.e. those with identical coordinates), retaining the best-supported model based on coverage and sequence identity. MAGIC founder Annotated assembly ENA V1.0 accession Annotated assembly ENA V1.1 accession Assembly FASTAs (same for both 1.0 and 1.1 annotations) GFF3 V1.1 Alchemy GCA_951799155.1 GCA_951799155.2 Triticum_aestivum_Alchemy_pseudomolecules_v1.fasta.gz Triticum_aestivum_Alchemy_pseudomolecules_v1.1_HC.gff3.gzTriticum_aestivum_Alchemy_pseudomolecules_v1.1_LC.gff3.gz Brompton GCA_964488245.1 GCA_964488245.2 Triticum_aestivum_Brompton_pseudomolecules_v1.fasta.gz Triticum_aestivum_Brompton_pseudomolecules_v1.1_HC.gff3.gzTriticum_aestivum_Brompton_pseudomolecules_v1.1_LC.gff3.gz Claire GCA_964494315.1 GCA_964494315.2 Triticum_aestivum_Claire_pseudomolecules_v1.fasta.gz Triticum_aestivum_Claire_pseudomolecules_v1.1_HC.gff3.gzTriticum_aestivum_Claire_pseudomolecules_v1.1_LC.gff3.gz Hereward GCA_964498875.1 GCA_964498875.2 Triticum_aestivum_Hereward_pseudomolecules_v1.fasta.gz Triticum_aestivum_Hereward_pseudomolecules_v1.1_HC.gff3.gzTriticum_aestivum_Hereward_pseudomolecules_v1.1_LC.gff3.gz Rialto GCA_964498365.1 GCA_964498365.2 Triticum_aestivum_Rialto_pseudomolecules_v1.fasta.gz Triticum_aestivum_Rialto_pseudomolecules_v1.1_HC.gff3.gzTriticum_aestivum_Rialto_pseudomolecules_v1.1_LC.gff3.gz Robigus GCA_964499505.1 GCA_964499505.2 Triticum_aestivum_Robigus_pseudomolecules_v1.fasta.gz Triticum_aestivum_Robigus_pseudomolecules_v1.1_HC.gff3.gzTriticum_aestivum_Robigus_pseudomolecules_v1.1_LC.gff3.gz Soissons GCA_964500185.1 GCA_964500185.2 Triticum_aestivum_Soissons_pseudomolecules_v1.fasta.gz Triticum_aestivum_Soissons_pseudomolecules_v1.1_HC.gff3.gzTriticum_aestivum_Soissons_pseudomolecules_v1.1_LC.gff3.gz Xi19 GCA_964506145.1 GCA_964506145.2 Triticum_aestivum_Xi19_pseudomolecules_v1.fasta.gz Triticum_aestivum_Xi19_pseudomolecules_v1.1_HC.gff3.gzTriticum_aestivum_Xi19_pseudomolecules_v1.1_LC.gff3.gz References Alaux, M., Rogers, J., Letellier, T., Flores, R., Alfama, F., Pommier, C., Mohellibi, N., Durand, S., Kimmel, E., Michotey, C., Guerche, C., Loaec, M., Lainé, M., Steinbach, D., Choulet, F., Rimbert, H., Leroy, P., Guilhot, N., Salse, J., Feuillet, C., … Quesneville, H. (2018). Linking the International Wheat Genome Sequencing Consortium bread wheat reference genome sequence to wheat genetic and phenomic data. Genome Biology, 19(1), 111. https://doi.org/10.1186/s13059-018-1491-4 Bushnell, B. (2014). BBMap: a fast, accurate, splice-aware aligner. Danecek, P., Bonfield, J. K., Liddle, J., Marshall, J., Ohan, V., Pollard, M. O., Whitwham, A., Keane, T., McCarthy, S. A., Davies, R. M., & Li, H. (2021). Twelve years of SAMtools and BCFtools. GigaScience, 10(2), giab008. https://doi.org/10.1093/gigascience/giab008 International Wheat Genome Sequencing Consortium (IWGSC) (2018). Shifting the limits in wheat research and breeding using a fully annotated reference genome. Science, 361(6403), eaar7191. https://doi.org/10.1126/science.aar7191 Jayakodi, M., Padmarasu, S., Haberer, G., Bonthala, V. S., Gundlach, H., Monat, C., Lux, T., Kamal, N., Lang, D., Himmelbach, A., Ens, J., Zhang, X. Q., Angessa, T. T., Zhou, G., Tan, C., Hill, C., Wang, P., Schreiber, M., Boston, L. B., Plott, C., … Stein, N. (2020). The barley pan-genome reveals the hidden legacy of mutation breeding. Nature, 588(7837), 284–289. https://doi.org/10.1038/s41586-020-2947-8 Li H. (2018). Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics, 34(18), 3094–3100. https://doi.org/10.1093/bioinformatics/bty191 Mackay, I. J., Bansept-Basler, P., Barber, T., Bentley, A. R., Cockram, J., Gosman, N., Greenland, A. J., Horsnell, R., Howells, R., O'Sullivan, D. M., Rose, G. A., & Howell, P. J. (2014). An eight-parent multiparent advanced generation inter-cross population for winter-sown wheat: creation, properties, and validation. G3, 4(9), 1603–1610. https://doi.org/10.1534/g3.114.012963 Monat, C., Padmarasu, S., Lux, T., Wicker, T., Gundlach, H., Himmelbach, A., Ens, J., Li, C., Muehlbauer, G. J., Schulman, A. H., Waugh, R., Braumann, I., Pozniak, C., Scholz, U., Mayer, K. F. X., Spannagl, M., Stein, N., & Mascher, M. (2019). TRITEX: chromosome-scale sequence assembly of Triticeae genomes with open-source tools. Genome Biology, 20(1), 284. https://doi.org/10.1186/s13059-019-1899-5 Padmarasu, S., Himmelbach, A., Mascher, M., & Stein, N. (2019). In Situ Hi-C for Plants: An Improved Method to Detect Long-Range Chromatin Interactions. Methods in Molecular Biology, 1933, 441–472. https://doi.org/10.1007/978-1-4939-9045-0_28 Percival-Alwyn, L., Barnes, I., Clark, M.D., Cockram, J., Coffey, M.P., Jones, S., Kersey, P.J., Kidner, C.A., Kosiol, C., Li, B., Marsh, W.A., Zhou, J., Caccamo, M., & Milne, I. (2024). UKCropDiversity‐HPC: A collaborative high-performance computing resource approach for sustainable agriculture and biodiversity conservation. Plants, People, Planet. https://doi.org/10.1002/ppp3.10607 Quinlan, A. R., & Hall, I. M. (2010). BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics, 26(6), 841–842. https://doi.org/10.1093/bioinformatics/btq033 Shumate, A., & Salzberg, S. L. (2021). Liftoff: accurate mapping of gene annotations. Bioinformatics, 37(12), 1639–1643. https://doi.org/10.1093/bioinformatics/btaa1016 Walkowiak, S., Gao, L., Monat, C., Haberer, G., Kassa, M. T., Brinton, J., Ramirez-Gonzalez, R. H., Kolodziej, M. C., Delorean, E., Thambugala, D., Klymiuk, V., Byrns, B., Gundlach, H., Bandi, V., Siri, J. N., Nilsen, K., Aquino, C., Himmelbach, A., Copetti, D., Ban, T., … Pozniak, C. J. (2020). Multiple wheat genomes reveal global variation in modern breeding. Nature, 588(7837), 277–283. https://doi.org/10.1038/s41586-020-2961-x

提供机构:
Zenodo
创建时间:
2025-12-11
二维码
社区交流群
二维码
科研交流群
商业服务