Natural variation in regulatory code revealed through Bayesian analysis of plant pan-genomes and pan-transcriptomes
收藏资源简介:
Understanding how non-coding DNA sequences control nearby gene expression is essential for synthetic biology and crop improvement. Current methods for identifying functional regulatory elements rely on expensive, specialized biochemical datasets and are typically limited to a single genotype, overlooking the genetic diversity present at the population level, especially in genes undergoing rapid evolution. Here we developed a computational tool that links natural sequence variation and gene expression variation across the pan-genome and pan-transcriptome to identify causative promoter sequences that directly regulate gene expression. Our tool employs a statistical framework that prioritizes causality over correlation, in contrast to most genome-wide association studies. Applying this tool to maize and soybean, two staple crops, we uncovered both known and novel regulatory elements often found in DNA structural variation that drive gene expression diversity. This approach provides a scalable and cost-effective way to map functional regulatory DNA in crops and efficiently utilizes natural variation from existing pangenomic datasets. It opens new avenues for crop engineering and the study of gene regulation in diverse plant species.



