遇见数据集

Data from "Identification of targets for introgression from rye (Secale ssp.) into wheat (Triticum aestivum) to increase arabinoxylan content in flour using a genome-wide association study"

收藏
Zenodo2026-07-27 更新2026-08-02 收录
官方服务:

资源简介:

This dataset contains the raw phenotypic and genotypic data relative to a collection of rye (Secale), as well as the input tables used to perform the genome-wide association study (GWAS) described in the publication associated with this dataset. The Secale collection used to generate this dataset comprised 312 rye accessions (Secale cereale subsp., S. strictum subsp., S. sylvestre, S. vavilovii) and five rye cultivars (“Helltop”, “Poseidon”, “Stannos”, “KWS Bono”, “Duiker Max”) which served as checks. The Secale collection was phenotypically characterised in the course of two field trials (S20/H21 and S21/H22) in terms of total arabinoxylan (TOT-AX), water-extractable arabinoxylan (WE-AX), and water-unextractable (WU-AX) content in flour, proportion of WE-AX tot TOT-AX content (WE/TOT-AX), and thousand kernel weight (TKW). Both field trials were laid out in an augmented design with six blocks. The number of rye genotypes (including the five check rye cultivars) phenotyped in trial S20/H21 was 267, and 235 in trial S21/H22. The number of rye genotypes (including the five check rye cultivars) phenotyped in trial S20/H21 as well as in trial S21/H22 was 208. Genotyping-by-sequencing (GBS) was used to genotypically characterise the Secale collection. DNA from pools of one to seven individual plants per accession of the Secale collection was used. A total of 294 genotypes (including the five check rye cultivars) were included in the GBS experiment on the Secale collection. A total of 37,907 unique GBS loci were obtained. A total of 49,498 SNPs were called from 5,976 of the obtained GBS loci. The minimum read depth required for SNP calling was set to 30. SNP calls were filtered based on call-rate ≥ 0.9, number of samples with two alleles ≥ 5, and reference allele frequency (RAF) between 0.05 and 0.95. Imputation of missing data was done by k-nearest neighbour algorithm (VIM R package, default settings; Kowarik & Templ, 2016). The best linear unbiased estimates (BLUE) of genotypic trait values and genotypic data relative to the 208 genotypes phenotyped in both field trials were used to perform a GWAS on all phenotypes measured using the BLINK model (GAPIT R package; Huang et al., 2019).

提供机构:
Zenodo
创建时间:
2026-07-09
二维码
社区交流群
二维码
科研交流群
商业服务