A Novel Statistical Framework for Big Data Prediction
A Novel Statistical Framework for Big Data Prediction
批准号:
1513408
负责人:
Shaw-Hwa Lo
金额:
$30.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-09-01 至 2019-08-31
中文摘要
全基因组关联研究的最新进展既增加了现有遗传数据的规模,也导致了导致各种疾病的重要遗传变异的确定。对这些遗传病的预测也变得至关重要。然而,对GWAS等大数据的预测并不是微不足道的。大数据预测中的一个关键障碍是识别(可能是少量的)变量集,当变量维度可能非常大时,这些变量集可以导致良好的预测。该项目探索了为什么一种常见的预测方法往往无法提供强大的预测率。将研究一种新的、基于交互和面向预测的方法来提取大数据中包含的隐藏信息。为了提高预测性,将开发一种新的标准来指导变量集的选择,优先考虑预测性而不是重要性,需要使用正确的预测率估计并制定基于预测性的标准来评估变量集。该项目提供了一个新的理论框架,通过表征什么有助于高度预测的变量集,并提供了一个新的标准来识别这些集的基础工作。在本研究项目的框架内,变量集具有理论(真实)可预测性水平,可以通过适当设计的基于样本的衡量标准来估计。这个框架是第一个寻求开发特定于可预测性标准的估计器的框架。此外,还将研究既包含边际影响又包含联合影响的方法,并研究可预测性的候选衡量标准。分析了四个真实的数据实例,以说明通过新方法找到的最终预测值与当前文献中的其他方法进行了比较。
英文摘要
Recent advances in genome-wide association studies (GWAS) have led to both an increase in the size of genetic data available and identification of important genetic variants responsible for a variety of diseases. Prediction for these genetic diseases has also become of paramount importance. However, prediction for big data such as GWAS is not trivial. A key obstacle in big data prediction is identifying (perhaps a small number of) variable sets that lead to good prediction when variable dimensionality can be extremely large. The project explores why a common approach towards prediction can often fail to deliver strong prediction rates. A novel, interaction-based and prediction-oriented approach to extracting hidden information contained in big data will be investigated. To improve prediction, a new criterion to guide the selection of variable sets will be developed.Prioritizing predictivity, not significance, requires using the correct estimates of prediction rates and developing predictivity-based criteria to evaluate variable sets. The project offers a novel theoretical framework by characterizing what makes for highly predictive variable sets, and providing fundamental work towards a new criterion to identify these sets. In the framework of this research project, variable sets have theoretical (true) levels of predictivity, which can be estimated with appropriately designed sample-based measures. This framework is the first that seeks to develop estimators specific to a criterion of predictivity. Additionally, methods that encompass both marginal and joint effects will be investigated, and a candidate measure of predictivity will be studied. Four real data examples are analyzed to illustrate how final predictors found via the new approach compare to other approaches in the current literature.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
BIGDATA: F: Statistical Foundation of Predictivity: A Novel Architecture for Big Data Learning
-
批准号:1741191
-
项目类别:Standard Grant
-
资助金额:$90.0万
-
财政年份:2018
-
负责人:Shaw-Hwa Lo
-
依托单位:
Collaborative Research: A General Framework for High Throughput Biological Learning: Theory Development and Applications
-
批准号:0714669
-
项目类别:Standard Grant
-
资助金额:$27.0万
-
财政年份:2007
-
负责人:Shaw-Hwa Lo
-
依托单位:
Statistical Analysis of Linkage/Association on Family-Based Studies in Human Genetics
-
批准号:0071930
-
项目类别:Continuing Grant
-
资助金额:$26.05万
-
财政年份:2000
-
负责人:Shaw-Hwa Lo
-
依托单位:
海外基金