Robust and Relevant Model Evaluation: Principles and Techniques for Handling Weak Prior Information and Contaminated Data
Robust and Relevant Model Evaluation: Principles and Techniques for Handling Weak Prior Information and Contaminated Data
批准号:
1209194
负责人:
Steven MacEachern
金额:
$32.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2012
资助国家:
美国
项目状态:
已结题
起止时间:
2012-09-01 至 2016-08-31
中文摘要
本研究关注创新模型比较和模型评估方法的发展,这些方法关注数据的最相关特征,并且对模型和数据的缺陷具有鲁棒性。 现代自动化数据收集技术的出现,以及廉价、近乎无限的存储容量的出现,提供了对以前无法想象的大量数据的访问。 并行开发复杂的模型,使人们能够联合收割机结合许多信息来源和计算策略和马力,使人们能够适应模型,似乎有助于近乎完美的决策。 然而,数据的丰富性使得数据污染带来的问题更加突出,模型的复杂性使得参数先验的确定更加困难。 处理数据污染和构建对缺乏先验信息具有鲁棒性的方法构成了基本的统计挑战。 在这个项目中,研究人员说明了目前领先的模型评估/模型比较方法的不足,然后提出了一套新的工具来缓解这些问题。 拟议的研究包括以下两个具体目标。 1. 在缺乏先验信息的情况下,发展可靠的贝叶斯模型比较方法。 使用不正确的无信息先验分布或模糊的正确先验分布的常见做法是有效的估计,但他们打破了贝叶斯假设检验问题,模型选择是敏感的先验分布的细节。 为了解决这一难题,研究人员提出了一种补救措施,校准贝叶斯因子。 校准贝叶斯因子不需要广泛的主观评价,产生的分析,更好地模仿贝叶斯因子的性能下的“合理的默认”之前,并广泛适用于各种各样的模型比较问题。 2. 在存在污染数据的情况下,开发用于模型评估和模型拟合的稳健方法。 受污染的数据有多种形式,包括可能来自记录错误或不相关人群的观察结果。 污染过程可能是不稳定的,这使得标准的统计建模不可行。 对于模型拟合,该项目开发并实施了限制似然,这导致了专注于数据最相关特征的估计策略,并且对“坏数据”具有鲁棒性。 对于模型评估,本项目开发了一种自适应损失(评分)交叉验证范式,通过稳定评估产生稳健的结果并产生上级有限样本性能。 模型评估和模型比较每天都在科学和企业决策环境中使用。 这些技术帮助研究人员判断哪种理论最能描述这种现象,帮助卫生专业人员确定哪些风险因素与疾病发生率有关,并帮助企业管理人员决定哪种商业策略可以增加销售或更好地留住客户。 然而,目前的模型评价和模型比较方法大多忽略了数据的不足或缺乏参数信息。 所提出的研究提供了强大的方法学工具,在这些困难的情况下,强大的模型偏好,模型评估和模型拟合。 它可以帮助各个领域的人们更好地从海量数据集中提取信息,从而优化他们的决策。 随着新方法的发展,健康研究、心理学实验和机器学习的具体应用将沿着进行。 一般的方法也适用于许多其他科学和技术领域,如基因组学,气候学和经济学,其中收集大量数据集和强大的模型评估是可取的。 调查人员完全有能力传播该项目的成果。 他们一直积极参与统计学和社会科学,工程/计算机科学和市场营销交叉的研究小组。 他们也是一个致力于提供和传播与保险业相关的研究的联合行业-大学中心的关键成员。 该项目的结果将通过调查人员与这些群体的互动传播到其他社区。
英文摘要
This research concerns the development of innovative model comparison and model evaluation methods that focus on the most relevant features of the data and that are robust to deficiencies of model and data. The advent of modern, automated technology for data collection and of cheap, near-boundless storage capacity provides access to a previously undreamt wealth of data. The parallel development of sophisticated models which allow one to combine many sources of information and the computational strategies and horsepower which allow one to fit the models would seem to facilitate near-perfect decision-making. However, the wealth of data aggravates the problems caused by data contamination, and the complexity of model aggravates the difficulty of specifying the prior on the parameters. Handling data contamination and constructing methods robust to the lack of prior information pose fundamental statistical challenges. In this project, the investigators illustrate the deficiency of the current leading model evaluation/model comparison methods, and then propose a set of new tools to alleviate the problems. The proposed research consists of the following two specific aims. 1. To develop reliable methods for Bayesian model comparison when prior information is lacking. The common practices of using an improper noninformative prior distribution or a vague proper prior distribution are effective in estimation, however they break down for Bayesian hypothesis testing problems where model choice is sensitive to details of the prior distribution. To tackle this difficulty, the investigators propose a remedy, the calibrated Bayes factor. The calibrated Bayes factor does not need extensive subjective evaluation, yields an analysis that better mimics the performance of the Bayes factor under a ''reasonable default'' prior, and is widely applicable in a large variety of model comparison problems. 2. To develop robust methods for model evaluation and model fitting in the presence of contaminated data. Contaminated data comes in many forms, including observations potentially from recording mistakes or from irrelevant populations. The contaminating process might be unstable, which makes standard statistical modeling infeasible. For model fitting, this project develops and implements restricted-likelihood, which leads to estimation strategies that focus on the most relevant features of the data and that are robust to ''bad data''. For model evaluation, this project develops an adaptive loss (scoring) paradigm for cross-validation, which produces robust results and yields superior finite-sample performance by stabilizing the evaluation. Model evaluation and model comparison are used on a daily basis in both scientific and corporate decision-making settings. These techniques help researchers judge which theory best describes the phenomenon, help health professionals identify which risk factors are related to disease incidence, and help corporate managers decide which business strategy results in increased sales or better customer retention. However, most of the current model evaluation and model comparison methods neglect deficiencies in data or suffer from the lack of parameter information. The proposed research provides powerful methodological tools for robust model preference, model evaluation, and model fitting in these difficult situations. It can help people in various fields better extract information from massive data sets, and thus optimize their decision making. Specific applications in health studies, psychological experiments and machine learning will proceed along with development of the new methodology. The general methodology is also applicable to many other scientific and technical areas, such as genomics, climatology, and economics, where large data sets are collected and robust model evaluation is desirable. The investigators are well-positioned to disseminate the project's results. They have been actively involved in research groups at the intersection of Statistics and the social sciences, engineering/computer science, and marketing. They are also key members of a joint industry-university center dedicated to provide and disseminate research relevant to the insurance industry. Results from this project will be spread to other communities through the investigators' interactions with these groups.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Statistical Inference under Subjective and Not-Fully-Quantifiable Information on Experimental Units
-
批准号:0605041
-
项目类别:Standard Grant
-
资助金额:$20.0万
-
财政年份:2006
-
负责人:Steven MacEachern
-
依托单位:
Nonparametric Bayesian Modelling
-
批准号:0072526
-
项目类别:Continuing Grant
-
资助金额:$8.8万
-
财政年份:2000
-
负责人:Steven MacEachern
-
依托单位:
海外基金