A common resequencing‐based genetic marker data set for global maize diversity

A common resequencing‐based genetic marker data set for global maize diversity
复制标题

全球玉米多样性的基于通用重测序的遗传标记数据集

DOI:
10.1111/tpj.16123
复制
发表时间:
2023
期刊:
The Plant Journal
影响因子:
--
通讯作者:
Schnable, James C.
Schnable, James C.
中科院分区:
--
文献类型:
--
作者:
Grzybowski, Marcin W.;Mural, Ravi V.;Xu, Gen;Turkus, Jonathan;Yang, Jinliang;Schnable, James C.

文献摘要

相似文献

玉米(Zea maysssp.mays)群体表现出广泛的遗传和表型多样性。随着测序成本的下降,越来越多的项目试图使用全基因组重测序策略来测量玉米群体之间和玉米群体内的遗传差异,识别数百万个分离的单核苷酸多态性(SNP)和插入/缺失(InDel)。与旧的基因分型策略(如微阵列和测序基因分型)不同,重测序原则上应该经常识别和评分常见的遗传变异。然而,在实践中,不同的项目经常采用不同的分析管道,经常采用不同的参考基因组组装,并一致地过滤研究人群中的次要等位基因频率。这限制了重新利用和重新混合不同项目产生的遗传多样性数据,以新方式解决新的生物问题的潜力。在这里,我们使用来自1276个先前发表的玉米样品和239个新重测序的玉米样品的重测序数据,以生成一个单一的统一标记集,其中包含约3.66亿个分离变体和约4600万个高置信度变体,这些变体在不同育种时期的作物野生近缘种、地方品种以及热带和温带品系中得分。我们证明,新的变体集提供了更大的能力,可以使用以前发表的性状数据集识别已知的因果开花时间基因,以及跟踪现代玉米全球分布中功能不同等位基因频率变化的潜力。
Maize (Zea maysssp.mays) populations exhibit vast ranges of genetic and phenotypic diversity. As sequencing costs have declined, an increasing number of projects have sought to measure genetic differences between and within maize populations using whole‐genome resequencing strategies, identifying millions of segregating single‐nucleotide polymorphisms (SNPs) and insertions/deletions (InDels). Unlike older genotyping strategies like microarrays and genotyping by sequencing, resequencing should, in principle, frequently identify and score common genetic variants. However, in practice, different projects frequently employ different analytical pipelines, often employ different reference genome assemblies and consistently filter for minor allele frequency within the study population. This constrains the potential to reuse and remix data on genetic diversity generated from different projects to address new biological questions in new ways. Here, we employ resequencing data from 1276 previously published maize samples and 239 newly resequenced maize samples to generate a single unified marker set of approximately 366 million segregating variants and approximately 46 million high‐confidence variants scored across crop wild relatives, landraces as well as tropical and temperate lines from different breeding eras. We demonstrate that the new variant set provides increased power to identify known causal flowering‐time genes using previously published trait data sets, as well as the potential to track changes in the frequency of functionally distinct alleles across the global distribution of modern maize.