遇见数据集

GBS data and morphological nut trait data from two populations of hazelnut (<em>Corylus</em> spp.)

收藏
NIAID Data Ecosystem2026-05-10 收录
官方服务:

资源简介:

The native, perennial shrub American hazelnut (Corylus americana) is cultivated in the Midwestern U.S. for its significant ecological benefits, as well as its high-value nut crop. Developing improved varieties for the Upper Midwest requires validated quantitative genetic approaches to utilizing genomic data in both American hazelnut and interspecific hybrids typical of breeding programs. In addition, high-throughput phenotyping methods are essential to the efficient and accurate screening of large breeding populations. This study reports novel advances in both of these domains. Two populations of hazelnuts, one composed of C. americana and one composed of C. americana x C. avellana hybrids, were phenotyped over the course of two years in two locations using a digital imagery-based method for quantifying morphological nut and kernel traits. This data was utilized to perform both composite interval mapping (CIM), using a recently released genetic map, and genomic prediction, using a newly available chromosome-scale reference genome for C. americana. Multiple QTL were detected for all traits analyzed, with an average total R2 of 52%. Marker-assisted genomic selection exhibited high prediction accuracy, with an average correlation coefficient between genotypic values and phenotypic observations of 0.78 across both environments. These results suggest that genomic prediction is a tenable method for improving genetic gain for highly polygenic traits in hazelnut breeding programs. Methods Phenotyping Morphological characteristics of in-shell hazelnuts and shelled kernels were used as the trait data for this study. This was collected primarily using an adapted version of the digital imagery acquisition and analysis pipeline reported by (Hameed et al. 2018). In brief, bushes were completely harvested by hand in August 2020 and August 2021. Harvested clusters were dried in the greenhouse, husked, and 30 in-shell nuts were then randomly sampled per bush. These nuts were arranged on a 6x5 grid with a QR code, and a Nikon 5600 DSLR camera tethered to a desktop computer was used to acquire a single image. The OpenCV Python library was used to isolate each in-shell nut and produce a binary mask by applying a fixed HSV threshold to each pixel. An ellipse was then fit to each binary mask, and the length of its major and minor axes was calculated by converting pixel length to physical distance using a scale bar embedded in each image. Circularity of the nut was calculated as the ratio of these two lengths. A bulk weight for the subsample was then obtained, and each nut was individually cracked. Each kernel was then returned to the grid, preserving the original arrangement of the nuts, and a second photo was acquired, and the same traits were calculated. Finally, a bulk weight for the kernels was also obtained. This allowed for both a volumetric and a gravimetric estimation of the percent kernel for each nut sampled. Python scripts for the image acquisition, processing, phenotyping, and file management are available at: https://github.com/shbrainard/hazelnut-phenotyping. Sequencing and genotyping Roughly 1 cm2~~ of leaf tissue was sampled from each bush in May 2021, immediately following budbreak. Tissue was sampled into 96-well Qiagen Collection Microtubes (Qiagen N.V., Venlo, The Netherlands) and lyophilized using a Labconco 18 L freeze dryer set to 0.004mBar for 72 hours. Freeze-dried tissue was then macerated. DNA extraction, quantification, library preparation, and sequencing were performed at the University of Wisconsin-Madison Biotechnology Center. Libraries were prepared for genotyping-by-sequencing using a double digestion with the restriction enzymes NsiI and BfaI, following the methodology described by Elshire et al. 2011). This combination was pre-selected based on an analysis of k-mer distributions of various enzyme digestions, where NsiI/BfaI was observed to maximize the k-mer diversity of the library. Illumina adapters and sample-specific barcodes were then annealed. Samples were multiplexed, and paired-end 150-bp sequence data were generated using an Illumina NovaSeq 6000, with an average of 10 million reads per sample. Trimming and demultiplexing of raw Illumina reads was performed using a custom Java application https://github.com/shbrainard/gbsTools. Reads were aligned to the C. americana genome for ‘Winkler’ (Brainard et al. 2024). Biallelic SNPs were utilized for the calculation of GEBVs, and were called using the TASSEL GBSv2 pipeline (Bradbury et al. 2007). SNPs were then filtered for missingness (<10% across all samples), linkage disequilibrium (r2 < 0.75), and allele depth (80th quantile of samples having a depth >8). Since the Wisconsin population of wild seedlings was unstructured, SNPs were also filtered to exclude sites with minor allele frequency < 0.05. All filtering was performed using bcftools (Danecek et al. 2021). Markers were then subset to only include those that were retained in both the Wisconsin and Minnesota populations, leaving 44,961 SNPs. For the construction of the genetic map, haplotype-based markers were called using Stacks 2 (Rochette et al. 2019), which can identify multi-allelic markers using the phased nature of multiple indels or SNPs that appear within a single 150-bp paired-end read. Because such markers cannot be directly filtered for depth, the parameter ‘gt-alpha’ was increased to 0.01 as a method for ensuring genotype quality. Markers were then filtered for linkage disequilibrium (r2 > 0.95) using bcftools. This generated a set of 78,079 markers, with an average of ~7,000 per chromosome. Linkage map construction and QTL analysis Since the Minnesota populations were constructed from controlled crosses between known hazelnut varieties, it was possible to build a genetic map by using the R package onemap (Margarido et al. 2007) (https://github.com/augusto-garcia/onemap). This map was previously reported in (Brainard et al. 2023). Briefly, markers called using Stacks 2 were first filtered to include only those of segregation types A1, A2, and B3.7 (following the notation of Wu et al. 2002), such that only markers with either three or four alleles remained. Next, markers for which more than 5% of all samples had no called genotype were removed, and two-point recombination frequencies were calculated for all possible phase configurations between all remaining markers using maximum likelihood. Maximum likelihood estimates were able to fully resolve the phase between pairs of markers, due to the fully informative segregation types that were utilized. A hierarchical clustering algorithm was utilized to construct linkage groups, and markers were ordered and phased within these groups by using a Hidden Markov Model with an error rate of 0.05. Recombination frequencies were finally converted to genetic distances using the Kosambi mapping function. Finally, this map was imported into the R package fullsibQTL (Gazaffi et al. 2020) (https://github.com/augusto-garcia/fullsibQTL), which was used to perform composite interval mapping. Calculation of GEBVs To calculate genomic-estimated breeding values from the biallelic SNP dataset described above, the R package StageWise (Endelman 2023) (https://github.com/jendelman/StageWise) was used to compute variance components and best linear unbiased predictors (BLUPs) of additive genetic value. This software is designed to perform a two-stage analysis, by first computing best linear unbiased estimators (BLUEs) for each genotype using a specified experimental design. Since the genotypes in both populations were comprised of unreplicated seedlings, a fixed-effects linear model was used to first compute each genotype’s BLUE.

创建时间:
2025-12-12
二维码
社区交流群
二维码
科研交流群
商业服务