Computational Methods for Phenotype Prediction to Assist Plant Breeding
Computational Methods for Phenotype Prediction to Assist Plant Breeding
批准号:
RGPIN-2021-04056
负责人:
Yan, Yan
金额:
$1.75万
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2021
资助国家:
加拿大
项目状态:
已结题
起止时间:
2021-01-01 至 2022-12-31
中文摘要
为了满足日益增长的人口对粮食的需求,作物育种效率需要大幅提高。我的研究项目的长期目标是开发一套人工智能工具,利用全基因组水平的信息(基因型)和环境因素来辅助植物育种。未来五年的短期目标是仅利用基因型数据预测作物的物理特性或性状(表型)。这将为育种者提供有效的性状选择和加快育种计划。表型预测具有挑战性,因为基因型数据中的特征数量明显多于样本数量。现有的方法通常需要非常大的计算资源,或者无法找到基因型和表型之间的联系。在本提案中,将开发多种策略来减少特征数量,增加样本量,提高预测精度。目的1是使用先进的采样算法减少基因型数据中的特征数量。我们将修改现有算法,使其适用于大型不平衡数据(如植物),而不会产生选择偏差。算法的性能将通过拟南芥、扁豆和小麦数据进行评估。所得到的特征将作为预测算法的可行输入,需要适度的计算资源。目标2是通过开发一种合成数据生成器来增加样本量,该生成器可以产生与植物数据具有相似特征的数据。将制定一套统计标准来衡量合成数据和真实数据之间的相似性。合成数据将使用实用的机器学习模型生成。数据生成器将为表型预测模型提供足够的训练数据(连同目标1的结果)。它还可以帮助减少建立大量植物基因型/表型数据的需要。目标3是结合目标1和目标2的结果,并使用深度学习(DL)模型预测植物表型。从我以前的工作中开发的模型已被证明对细菌数据(具有少量特征)是有效的,并且将被修改以适合植物数据。此外,将在DL模型中添加一个解释层来解释结果。领域专家可以读取可解释的信息并验证预测。影响将体现在三个方面:1。它将提供一套相似性测量的统计标准,从而促进数据比较标准的发展。2. 它将通过提示与植物表型可靠相关的基因组特征来加快和加强植物育种中的选择。3. 通过开发可推广到其他领域的深度学习模型,它将有助于解决具有大特征小样本数据的问题。
英文摘要
To meet the food demands of an increasing population, crop breeding efficiency needs to be substantially improved. The long-term goal of my research program is to develop a suite of Artificial Intelligence tools to assist plant breeding by utilizing the whole genome level information (genotype) as well as environmental factors. The short-term goals in the next five years are to predict crops' physical properties or traits (phenotype) using genotype data only. This will provide breeders with effective trait selections and accelerate their breeding programs. Phenotype prediction is challenging because the number of features in genotype data is significantly more than the number of samples. Existing methods usually require extraordinarily large computational resources or fail to find the linkage between genotype and phenotype. In this proposal, multiple strategies will be developed to reduce the number of features, increase the sample size, and improve the prediction accuracy. Objective 1 is to reduce the number of features in the genotype data using advanced sampling algorithms. We will modify existing algorithms to make them suitable for large imbalanced data (like the plant) without creating selection bias. The performance of the algorithms will be evaluated by Arabidopsis thaliana, lentil, and wheat data. The resulting features will serve as a feasible input to a prediction algorithm with modest computational resources required. Objective 2 is to increase the sample size by developing a synthetic data generator that can produce data with similar characteristics to plant data. A set of statistical criteria will be developed to measure the similarities between synthetic and real data. The synthetic data will be generated using a practical machine learning model. The data generator will provide sufficient training data (together with results from Objective 1) for the phenotype prediction model. It can also help reduce the need to establish extremely large collections of plant genotype/phenotype data. Objective 3 is to incorporate results from Objectives 1&2 and predict plant phenotypes using a Deep Learning (DL) model. The model developed from my previous work has proven to be effective on bacteria data (which has a small number of features) and will be modified to fit for the plant data. Further, an interpretation layer will be added to the DL model to explain the results. Domain experts can read into the interpretable information and validate the predictions. The impact will be in three aspects: 1. It will advance the development of data comparison standards, by providing a set of statistical criteria for similarity measurement. 2. It will speed up and enhance the selection in plant breeding, by suggesting genomic characteristics that are reliably associated with plant phenotypes. 3. It will contribute to solving the problems that have large-feature-small-sample data, by developing DL models that can be generalized to other fields.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Computational Methods for Phenotype Prediction to Assist Plant Breeding
-
批准号:RGPIN-2021-04056
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.75万
-
财政年份:2022
-
负责人:Yan, Yan
-
依托单位:
Computational Methods for Phenotype Prediction to Assist Plant Breeding
-
批准号:DGECR-2021-00348
-
项目类别:Discovery Launch Supplement
-
资助金额:$0.91万
-
财政年份:2021
-
负责人:Yan, Yan
-
依托单位:
国内基金
海外基金
Computational Methods for Analyzing Toponome Data
-
批准号:60601030
-
项目类别:青年科学基金项目
-
资助金额:17.0万元
-
批准年份:2006
-
负责人:Axel Mosig
-
依托单位: