课题基金 / 基金详情

Martingale Control of mFDR in Variable Selection

Martingale Control of mFDR in Variable Selection
变量选择中 mFDR 的鞅控制
批准号:
1106743
负责人:
Robert Stine
金额:
$32.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2011
资助国家:
美国
项目状态:
已结题
起止时间:
2011-07-01 至 2014-06-30

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
这个项目的研究人员开发了在建立统计模型时控制从多个来源选择预测特征的方法。伪变量数量的鞅表示提供了潜在的理论支持。这个鞅定义了一个框架,用于测试可能无限的假设序列。这种表示导致流特征选择的方法,控制错误发现的预期数量(mFDR)。在这个项目中开发的扩展概括了研究人员先前的工作,将他们的结果扩展到潜在特征的多个流,同时保持鞅表示。鉴于作者之前的工作是在高噪声,低信号设置中,其中很少有特征是可预测的(近黑设置),该提案的进展将他们的方法推向了具有更高信噪比的许多预测特征的问题。该建议设想用与模型拟合优度直接相关的鞅代替原始鞅。研究人员计划使用这个修订的鞅来证明一个基于拍卖的系统,它结合了几个特征源,满足mFDR条件。研究人员开发了新的方法来建立预测统计模型,结合并从多个信息来源学习。预测统计模型是根据数据构建的经验规则,它根据其他特征的值预测观测的特定特征,即响应。构建这些模型的挑战在于确定能够产生预测性见解的特征。虽然越来越多的数据是统计模型的重要输入,但大量特征的存在导致了过度拟合的问题。当人们将特征之间的随机巧合与可重复的模式混淆时,就会出现过度拟合。现代数据挖掘产生了如此多的特征,以至于很难区分真实的关联和虚构的关联。研究人员提出了一个系统,在一个共同的建模范式的背景下,使这些区别。作为一个实际的测试平台,研究者将使用回归分析来分析经典的计算语言问题,这是应用统计学的主力方法。考虑到语言学的经验范围,回归模型的任何缺陷都会突出。这将鼓励回归的创新,在与语言学中手工制作的方法竞争的同时保持其简单性。这些创新应该扩展到其他应用,包括功能磁共振成像、遗传学和更一般的数据挖掘。
英文摘要
The investigators in this project develop methods that control the selection of predictive features from multiple sources when building statistical models. A martingale representation of the number of spurious variables provides the underlying theoretic support. This martingale defines a framework for testing a possibly infinite sequence of hypotheses. This representation leads to methods for streaming feature selection that control the expected number of false discoveries (mFDR). Extensions to be developed in this project generalize prior work of the investigators, extending their results to multiple streams of potential features while maintaining the martingale representation. Whereas the previous work of the authors was in the high-noise, low-signal setting in which few features are predictive (the nearly black setting), advances in this proposal push their methods into problems characterized by many predictive features with higher signal-to-noise ratios. This proposal envisions replacing the original martingale by one directly related to the goodness of fit of the model. The investigators plan to use this revised martingale to show that an auction-based system that combines several sources of features satisfies the mFDR condition.The investigators develop novel methods for building predictive statistical models that combine and learn from multiple sources of information. A predictive statistical model is an empirical rule constructed from data that predicts a specific characteristic of observations, the response, based on the values of other characteristics. The challenge of building these models is to identify characteristics that yield predictive insights. While ever larger amounts of data are an essential input to a statistical model, the presence of vast numbers of characteristics lead to the problem of over-fitting. Over-fitting occurs when one confuses a random coincidence among characteristics with a reproducible pattern. Modern data mining produces such a plethora of characteristics that it becomes difficult to distinguish real from imaginary associations. The investigators propose a system that makes these distinctions in the context of a common modeling paradigm. As a practical testbed, the investigators will analyze classic computational linguistic problems using regression analysis, the workhorse method of applied statistics. Given the extent of experience in linguistics, any deficiencies of a regression model will stand out. This will encourage innovations in regression that maintain their simplicity while competing with handcrafted methods in linguistics. These innovations should extend to other applications including fMRI, genetics, and more general data mining.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
Cortical control of internal state in the insular cortex-claustrum region