Robust and Relevant Model Evaluation: Principles and Techniques for Handling Weak Prior Information and Contaminated Data
Robust and Relevant Model Evaluation: Principles and Techniques for Handling Weak Prior Information and Contaminated Data
批准号:
1209194
负责人:
Steven MacEachern
金额:
$32.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2012
资助国家:
美国
项目状态:
已结题
起止时间:
2012-09-01 至 2016-08-31
中文摘要
这项研究涉及到创新的模型比较和模型评估方法的发展,这些方法侧重于数据的最相关特征,并且对模型和数据的缺陷具有健壮性。现代自动化数据收集技术的出现,以及廉价、近乎无限的存储容量的出现,为人们提供了访问以前做梦也想不到的海量数据的途径。同时开发复杂的模型,使人们能够结合许多信息源,以及使人们能够匹配模型的计算策略和马力,似乎有助于做出近乎完美的决策。然而,数据的丰富加剧了数据污染带来的问题,模型的复杂性增加了参数先验的确定难度。处理数据污染和构建对缺乏先验信息的健壮方法构成了基本的统计挑战。在这个项目中,研究人员阐明了目前领先的模型评估/模型比较方法的不足,并提出了一套新的工具来缓解这些问题。拟议的研究包括以下两个具体目标。1.在缺乏先验信息的情况下,发展可靠的贝叶斯模型比较方法。使用不适当的无信息先验分布或模糊适当的先验分布的常见做法在估计中是有效的,但对于模型选择对先验分布的细节敏感的贝叶斯假设检验问题来说,它们是失败的。为了解决这一困难,研究人员提出了一种补救措施,即校准贝叶斯因子。校准的贝叶斯因子不需要广泛的主观评估,所产生的分析更好地模拟了贝叶斯因子在“合理违约”先验情况下的表现,并且广泛适用于各种模型比较问题。2.发展在污染数据存在的情况下进行模型评估和模型拟合的稳健方法。受污染的数据有多种形式,包括可能来自记录错误或来自无关人群的观察。污染过程可能是不稳定的,这使得标准的统计建模不可行。对于模型拟合,这个项目开发并实现了受限似然,这导致了专注于数据的最相关特征的估计策略,并且对“坏数据”具有健壮性。对于模型评估,该项目开发了一种用于交叉验证的自适应损失(评分)范例,它通过稳定评估来产生稳健的结果并产生优越的有限样本性能。在科学和公司决策环境中,每天都会使用模型评估和模型比较。这些技术帮助研究人员判断哪种理论最能描述这种现象,帮助卫生专业人员确定哪些风险因素与疾病发病率有关,并帮助公司经理决定哪种商业战略可以增加销售额或更好地留住客户。然而,目前的模型评价和模型比较方法大多忽视了数据的不足或缺乏参数信息。所提出的研究为稳健的模型偏好、模型评估以及在这些困难情况下的模型拟合提供了强大的方法工具。它可以帮助各个领域的人们更好地从海量数据集中提取信息,从而优化他们的决策。随着新方法的发展,在健康研究、心理学实验和机器学习方面的具体应用将继续进行。一般方法也适用于许多其他科学和技术领域,如基因组学、气候学和经济学,这些领域收集了大量数据,需要进行稳健的模型评估。调查人员处于有利地位,可以公布该项目的结果。他们积极参与了统计学和社会科学、工程/计算机科学和市场营销学交叉学科的研究小组。他们也是一个产学联合中心的关键成员,该中心致力于提供和传播与保险业相关的研究。该项目的结果将通过调查人员与这些团体的互动传播到其他社区。
英文摘要
This research concerns the development of innovative model comparison and model evaluation methods that focus on the most relevant features of the data and that are robust to deficiencies of model and data. The advent of modern, automated technology for data collection and of cheap, near-boundless storage capacity provides access to a previously undreamt wealth of data. The parallel development of sophisticated models which allow one to combine many sources of information and the computational strategies and horsepower which allow one to fit the models would seem to facilitate near-perfect decision-making. However, the wealth of data aggravates the problems caused by data contamination, and the complexity of model aggravates the difficulty of specifying the prior on the parameters. Handling data contamination and constructing methods robust to the lack of prior information pose fundamental statistical challenges. In this project, the investigators illustrate the deficiency of the current leading model evaluation/model comparison methods, and then propose a set of new tools to alleviate the problems. The proposed research consists of the following two specific aims. 1. To develop reliable methods for Bayesian model comparison when prior information is lacking. The common practices of using an improper noninformative prior distribution or a vague proper prior distribution are effective in estimation, however they break down for Bayesian hypothesis testing problems where model choice is sensitive to details of the prior distribution. To tackle this difficulty, the investigators propose a remedy, the calibrated Bayes factor. The calibrated Bayes factor does not need extensive subjective evaluation, yields an analysis that better mimics the performance of the Bayes factor under a ''reasonable default'' prior, and is widely applicable in a large variety of model comparison problems. 2. To develop robust methods for model evaluation and model fitting in the presence of contaminated data. Contaminated data comes in many forms, including observations potentially from recording mistakes or from irrelevant populations. The contaminating process might be unstable, which makes standard statistical modeling infeasible. For model fitting, this project develops and implements restricted-likelihood, which leads to estimation strategies that focus on the most relevant features of the data and that are robust to ''bad data''. For model evaluation, this project develops an adaptive loss (scoring) paradigm for cross-validation, which produces robust results and yields superior finite-sample performance by stabilizing the evaluation. Model evaluation and model comparison are used on a daily basis in both scientific and corporate decision-making settings. These techniques help researchers judge which theory best describes the phenomenon, help health professionals identify which risk factors are related to disease incidence, and help corporate managers decide which business strategy results in increased sales or better customer retention. However, most of the current model evaluation and model comparison methods neglect deficiencies in data or suffer from the lack of parameter information. The proposed research provides powerful methodological tools for robust model preference, model evaluation, and model fitting in these difficult situations. It can help people in various fields better extract information from massive data sets, and thus optimize their decision making. Specific applications in health studies, psychological experiments and machine learning will proceed along with development of the new methodology. The general methodology is also applicable to many other scientific and technical areas, such as genomics, climatology, and economics, where large data sets are collected and robust model evaluation is desirable. The investigators are well-positioned to disseminate the project's results. They have been actively involved in research groups at the intersection of Statistics and the social sciences, engineering/computer science, and marketing. They are also key members of a joint industry-university center dedicated to provide and disseminate research relevant to the insurance industry. Results from this project will be spread to other communities through the investigators' interactions with these groups.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Statistical Inference under Subjective and Not-Fully-Quantifiable Information on Experimental Units
-
批准号:0605041
-
项目类别:Standard Grant
-
资助金额:$20.0万
-
财政年份:2006
-
负责人:Steven MacEachern
-
依托单位:
Nonparametric Bayesian Modelling
-
批准号:0072526
-
项目类别:Continuing Grant
-
资助金额:$8.8万
-
财政年份:2000
-
负责人:Steven MacEachern
-
依托单位:
海外基金