课题基金 / 基金详情

Predictive Methods for Analyzing High-throughput Data and Spatial-Temporal Data

Predictive Methods for Analyzing High-throughput Data and Spatial-Temporal Data
分析高通量数据和时空数据的预测方法
批准号:
RGPIN-2019-07020
负责人:
Li, Longhai
金额:
$1.46万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2020
资助国家:
加拿大
项目状态:
已结题
起止时间:
2020-01-01 至 2021-12-31

项目摘要

项目成果

Li, Longhai的其他基金

相似基金

相关文献

中文摘要
翻译
高通量测序生物技术的加速发展使得收集高维分子水平的图谱变得负担得起,例如基因表达,这些在本提案中被称为特征。识别与表型相关的特征是非常有意义的(例如。癌症状况、健康障碍)。许多研究人员主张应用统计学习方法对高通量数据进行预测分析。预测分析结果可以在许多方面使用。例如,它们可以用来诊断人类疾病,预测对药物的反应(个性化药物);它们可以用来选择最优的基因子集,供动植物育种者进一步实验;从良好的预测模型中提取的特征子集可以帮助揭示表型的生物学机制。不幸的是,即使使用非常简单的模型,高维数据也会在预测分析中造成巨大的过度拟合。发现错误预测特征/模式的机会极高。因此,在搜索更具预测性的特征时,如何对抗预测分析中的错误发现是一项具有挑战性的工作。我的研究成果将包括诚实地衡量选定特征的预测性(如错误率,AUC)的新工具,以及识别真正的预测性特征和建立更清晰的表型预测模型的新工具。我还将在与人类健康和食品安全有关的各种科学问题上使用特定的高通量数据集进行预测分析,这将为这些领域带来新的科学发现和新的解决方案。 在科学上,一个理论是通过对未来的观测进行预测来检验的。观测和预测之间的巨大差异表明,该理论是不正确的或有缺陷的。同样,查看样本外预测是比较和检查统计模型的拟合优度(GOF)的直接方法。今天,越来越复杂的模型被提出用于各种相关的数据,如时间、空间和重复测量数据。需要更广泛适用的预测方法来比较和检验这种复杂的模型。我将致力于改进具有相关性结构的数据集的广义线性混合模型(GLMM)的预测模型比较和检验方法,并发布R附加程序包,以方便GLMM的比较和检验。我的研究成果将包括评估具有相关随机效应的复杂贝叶斯/非贝叶斯模型的新工具。这些新的模型评估工具对于流行病学、生态学和环境科学的研究人员来说将是必不可少的。对这些领域的数据集进行改进的建模将导致更可靠的数据分析结论,这对经济-社会问题的政策制定具有重要影响。
英文摘要
The accelerated development of high-throughput sequencing biotechnologies has made it affordable to collect high-dimensional molecular-level profiles, such as gene expression, which are called features in this proposal. It is of great interest to identify relevant features associated with a phenotype (eg. cancer status, health disorder). Many researchers have advocated to apply statistical learning methods to perform predictive analysis for high-throughput data. Predictive analysis results can be used in many ways. For example, they can be used to diagnose human diseases, to predict response to a medicine (personalized medicine); they can be used to choose an optimal gene subset for further experiments by plant/animal breeders; the subset of features extracted from good predictive models can facilitate the uncovering of the biological mechanism for a phenotype. Unfortunately, the high-dimensionality causes enormous overfitting in predictive analysis even with very simple models. The chance of finding false predictive features/patterns is extremely high. Therefore, it is challenging to fight against false discovery in predictive analysis when searching for more predictive features. My research outcomes will include new tools for honestly measuring predictivity (such as error rate, AUC) of selected features, and new tools for identifying truly predictive features and for building sharper predictive models for phenotypes. I will also practice predictive analysis with specific high-throughput datasets in a variety of scientific problems related to human health and food security, which will lead to new scientific discoveries and new solutions for these areas. In science, a theory is tested by performing predictions for observations in the future. Significant discrepancies between observations and predictions suggest that the theory is incorrect or flawed. Similarly, looking at out-of-sample predictions is a straightforward method for comparing and checking goodness-of-fit (GOF) of statistical models. Today, increasingly complex models are being proposed for a variety of correlated data such as, temporal, spatial, and repeated measurements data. More widely applicable predictive methods for comparing and checking such complex models are demanded. I will work to improve predictive model comparison and checking methods for generalized linear mixed models (GLMM) for datasets with correlation structure, and to release R add-on packages to facilitate the comparison and checking of GLMM. My research outcomes will include new tools for evaluating complex Bayesian/non-Bayesian models with correlated random effects. These new model evaluation tools will be essential for researchers in epidemiology, ecology, and environmental sciences. Improved modelling of the datasets from these areas will lead to more solid data analysis conclusions, which have essential impact on policy making in economic-social problems.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Predictive Methods for Analyzing High-throughput Data and Spatial-Temporal Data
  • 批准号:
    RGPIN-2019-07020
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.46万
  • 财政年份:
    2022
  • 负责人:
    Li, Longhai
  • 依托单位:
Predictive Methods for Analyzing High-throughput Data and Spatial-Temporal Data
  • 批准号:
    RGPIN-2019-07020
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.46万
  • 财政年份:
    2021
  • 负责人:
    Li, Longhai
  • 依托单位:
Predictive Methods for Analyzing High-throughput Data and Spatial-Temporal Data
  • 批准号:
    RGPIN-2019-07020
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.46万
  • 财政年份:
    2019
  • 负责人:
    Li, Longhai
  • 依托单位:
Bayesian Methods for High-dimensional and Correlated Data
  • 批准号:
    RGPIN-2014-05010
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.02万
  • 财政年份:
    2018
  • 负责人:
    Li, Longhai
  • 依托单位:
国内基金
海外基金
Computational Methods for Analyzing Toponome Data