课题基金 / 基金详情

A Novel Statistical Framework for Big Data Prediction

A Novel Statistical Framework for Big Data Prediction
用于大数据预测的新型统计框架
批准号:
1513408
负责人:
Shaw-Hwa Lo
金额:
$30.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-09-01 至 2019-08-31

项目摘要

项目成果

Shaw-Hwa Lo的其他基金

相似基金

相关文献

中文摘要
翻译
全基因组关联研究(GWAS)的最新进展既增加了现有遗传数据的规模,也确定了导致各种疾病的重要遗传变异。对这些遗传疾病的预测也变得极为重要。然而,对像GWAS这样的大数据的预测并非微不足道。大数据预测的一个关键障碍是识别(可能是少数)变量集,当变量维度可能非常大时,这些变量集可以导致良好的预测。该项目探讨了为什么一种常见的预测方法往往不能提供高预测率。将研究一种新的、基于交互和面向预测的方法来提取大数据中包含的隐藏信息。为了提高预测精度,将提出一种新的准则来指导变量集的选择。优先考虑可预测性,而不是显著性,需要使用预测率的正确估计,并开发基于可预测性的标准来评估变量集。该项目通过描述高预测性变量集的特征提供了一个新的理论框架,并为识别这些集的新标准提供了基础工作。在本研究项目的框架中,变量集具有理论(真实)预测水平,可以通过适当设计的基于样本的度量来估计。这个框架是第一个寻求开发特定于可预测性标准的估计器的框架。此外,将研究包括边际效应和联合效应的方法,并研究可预测性的候选测量方法。分析了四个真实数据示例,以说明通过新方法发现的最终预测因子与当前文献中的其他方法相比如何。
英文摘要
Recent advances in genome-wide association studies (GWAS) have led to both an increase in the size of genetic data available and identification of important genetic variants responsible for a variety of diseases. Prediction for these genetic diseases has also become of paramount importance. However, prediction for big data such as GWAS is not trivial. A key obstacle in big data prediction is identifying (perhaps a small number of) variable sets that lead to good prediction when variable dimensionality can be extremely large. The project explores why a common approach towards prediction can often fail to deliver strong prediction rates. A novel, interaction-based and prediction-oriented approach to extracting hidden information contained in big data will be investigated. To improve prediction, a new criterion to guide the selection of variable sets will be developed.Prioritizing predictivity, not significance, requires using the correct estimates of prediction rates and developing predictivity-based criteria to evaluate variable sets. The project offers a novel theoretical framework by characterizing what makes for highly predictive variable sets, and providing fundamental work towards a new criterion to identify these sets. In the framework of this research project, variable sets have theoretical (true) levels of predictivity, which can be estimated with appropriately designed sample-based measures. This framework is the first that seeks to develop estimators specific to a criterion of predictivity. Additionally, methods that encompass both marginal and joint effects will be investigated, and a candidate measure of predictivity will be studied. Four real data examples are analyzed to illustrate how final predictors found via the new approach compare to other approaches in the current literature.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
BIGDATA: F: Statistical Foundation of Predictivity: A Novel Architecture for Big Data Learning
  • 批准号:
    1741191
  • 项目类别:
    Standard Grant
  • 资助金额:
    $90.0万
  • 财政年份:
    2018
  • 负责人:
    Shaw-Hwa Lo
  • 依托单位:
Collaborative Research: A General Framework for High Throughput Biological Learning: Theory Development and Applications
  • 批准号:
    0714669
  • 项目类别:
    Standard Grant
  • 资助金额:
    $27.0万
  • 财政年份:
    2007
  • 负责人:
    Shaw-Hwa Lo
  • 依托单位:
Statistical Analysis of Linkage/Association on Family-Based Studies in Human Genetics
  • 批准号:
    0071930
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $26.05万
  • 财政年份:
    2000
  • 负责人:
    Shaw-Hwa Lo
  • 依托单位:
海外基金