ABI Innovation: Improved modeling of marker-trait associations in polypoid and diploid organisms using genotyping-by-sequencing with genotype uncertainty
ABI Innovation: Improved modeling of marker-trait associations in polypoid and diploid organisms using genotyping-by-sequencing with genotype uncertainty
批准号:
1661490
负责人:
Lindsay Clark
金额:
$66.95万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-07-01 至 2022-06-30
中文摘要
许多重要的经济作物每条染色体都有两个以上的拷贝--它们的基因组中每个基因都有两个以上的等位基因,其重要性状的基因类型应该反映多倍体的水平。在育种项目中,基因类型被用来预测开花时间、种子重量或产油量等性状的表型。这些性状的标记基因类型目前是在定向测序实验中确定的,称为逐个测序的基因型(GBS)。GBS极大地降低了检测遗传变异和确定个体基因类型的成本,从而使农业和生态领域更容易进行高通量基因分型。然而,使用GBS确定的标记基因型往往是不正确的,因为由于随机的机会,一些等位基因在一些个体中没有得到测序。多倍体使基因多义性的问题进一步复杂化,因为基因类型不再简单地被归类为纯合或杂合,而是由它们所拥有的每个等位基因的拷贝数来定义。该项目旨在开发方法和软件来量化DNA标记基因分型(GBS)确定的不确定性,特别是在多倍体物种中,并将这种不确定性纳入将基因型别与表型相关的分析中。鉴于许多具有重要经济意义的作物是多倍体的,该项目带来的基因鉴定改进将特别有利于通过标记辅助植物育种确保世界粮食、燃料和纤维的供应。该项目将提高识别与表型差异相关的标记等位基因的敏感性,并改进从GBS数据中预测表型的方法。更广泛地说,它将推动该领域的范式转变,将基因类型视为概率分布,而不是确切已知的值。此外,在这个项目的过程中,将制作讲授线性代数以及用Python和R编写计算机程序的YouTube视频,使生物科学的学生更容易接触到这些概念。本项目的目标是(1)创建一种算法,尽可能准确地从GBS数据中估计二倍体和多倍体的基因概率,以及(2)创建充分利用基因概率的全基因组关联研究(GWAS)和基因组选择(GS)方法。将开发一种迭代算法,该算法将在给定的基因座上生成每个个体是每个可能的基因的概率分布。该算法将使用两个或更多等位基因中每一个的读取深度,并将对多个生物学和技术参数进行建模,包括等位基因频率、种群结构、近亲繁殖、连锁不平衡、遗传模式、差异扩增和样本污染。该算法将在一个公开可用的R包中实施,该包将与现有的GBS生物信息学管道集成,并将输出适合于下游分析的多种格式。将开发新的GWAS和GS方法,在利用基因概率分布的同时对加性和显性效应进行建模,并将以现有的GWAS和GS软件为基础。生物能源牧草芒属的二倍体和四倍体群体将被用于验证和改进新的方法。已经或正在对这些人群进行GBS和表型鉴定。该项目产生的软件将托管在https://github.com/lvclark/polyRAD和https://github.com/lvclark/GAPITdom上。
英文摘要
Many economically important crop plants have more than two copies of each chromosome - they have more than 2 alleles of each gene in their genome and the genotypes for their important traits should reflect the level of polyploidy. Genotypes are used to predict the phenotype in breeding programs, for traits like flowering time, seed weight or oil yield. Marker genotypes for such traits are currently determined in targeted sequencing experiments, termed genotype-by-sequencing (GBS). GBS has drastically reduced the cost of detecting genetic variants and determining the genotypes of individuals, thus making high-throughput genotyping more accessible within the fields of agriculture and ecology. However, marker genotypes determined using GBS are frequently incorrect because, due to random chance, some alleles do not get sequenced in some individuals. Polyploidy further complicates the issue of genotype ambiguity because genotypes can no longer simply be classified as homozygous or heterozygous, but instead are defined by the number of copies of each allele that they possess. This project aims to develop methodology and software for quantifying uncertainty in DNA marker genotypes determined by genotyping-by-sequencing (GBS), particularly in polyploid species, and for incorporating that uncertainty into analyses that relate genotype to phenotype. Given that many economically-important crops are polyploid, genotyping improvements that result from this project will be especially beneficial for securing the world's supply of food, fuel, and fiber through marker-assisted plant breeding. This project will result in increased sensitivity for identifying marker alleles that are associated with phenotypic differences, as well as improved prediction of phenotypes from GBS data. More broadly, it will promote a paradigm shift in the field, treating genotypes as probability distributions rather than values that are known with certainty. Additionally, YouTube videos teaching linear algebra as well as computer programming in Python and R will be created during the course of this project, making these concepts more accessible for students in the biological sciences.The objectives of this project are to (1) create an algorithm that estimates genotype probabilities in diploids and polyploids as accurately as possible from GBS data and (2) create genome-wide association study (GWAS) and genomic selection (GS) methodologies that fully utilize genotype probabilities. An iterative algorithm will be developed that will generate a probability distribution of each individual being each possible genotype at a given locus. The algorithm will use read depth at each of two or more alleles and will model multiple biological and technical parameters, including allele frequencies, population structure, inbreeding, linkage disequilibrium, inheritance mode, differential amplification, and sample contamination. The algorithm will be implemented in a publicly-available R package that will integrate with existing GBS bioinformatics pipelines and will output multiple formats suitable for downstream analysis. New GWAS and GS methods will be developed that model both additive and dominance effects while utilizing genotype probability distributions, and will build upon existing software for GWAS and GS. Diploid and tetraploid populations of the bioenergy grass Miscanthus will be used for validating and improving the new methodologies. GBS and phenotyping have already been performed or are underway on these populations. Software produced as a result of this project will be hosted at https://github.com/lvclark/polyRAD and https://github.com/lvclark/GAPITdom .
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
DOI:
10.1534/g3.118.200913
发表时间:
2019-03-01
期刊:
G3-GENES GENOMES GENETICS
影响因子:
2.6
作者:
[Clark, Lindsay V., Lipka, Alexander E., Sacks, Erik J.]
通讯作者:
Sacks, Erik J.
海外基金