Methods for Dealing with Misspecification in Bayesian Experimental Design
Methods for Dealing with Misspecification in Bayesian Experimental Design
批准号:
2740638
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2022
资助国家:
英国
项目状态:
未结题
起止时间:
2022 至 --
中文摘要
统计学和机器学习方面的许多研究都集中在分析已经获得的数据的方法上,但如何从一开始就最好地收集数据的问题却没有得到充分的探讨。收集数据可能是昂贵的,因此从业者受到他们可以收集的数据量的限制。如果不小心这样做,就会产生低质量的数据——可能导致不准确的结果和不正确的结论,无论一个人的分析工具有多先进。因此,在分析之前努力收集高质量的数据是至关重要的。实验设计研究旨在解决这一问题,为实践者提供收集信息数据的方法,这些方法将导致可靠的结果和强有力的结论。为了概述贝叶斯实验设计(BED),我们考虑以下设置:在一个区域内有几个信标,每个信标发出一个信号,从业者希望定位信标。数据收集过程包括从业者选择探测信号的位置,然后记录这些位置的信号强度。有了无限的资源,从业者将能够完美地定位信标,但在实践中,他们只能进行有限数量的实验。然后BED的工作如下:在收集任何数据之前,从业者将首先根据信标的未知位置形成给定位置信号强度的统计模型,他们将指定他们对信标位置的先验信念,并在这些位置上具有先验分布。然后,BED程序可以为从业者提供探测信号的最佳位置——这里的“最佳”定义为将导致有关信标位置的信息增加最多的位置。上述情况可以很容易地推广到其他设置。BED在理论上是合理的,在实践中也表现良好,然而,如果我们的数据统计模型被错误地指定,也就是说,如果真实的数据生成过程与我们指定的模型不同,它就会崩溃。贝叶斯统计方法总是容易受到模型规格错误的影响,但不幸的是,这个事实对于实验设计来说尤其有问题,因为我们不仅要使用我们的模型来分析数据,还要首先收集数据。在最坏的情况下,有些模型的最佳行动方案是在完全相同的地方选择所有的设计,而不管你观察到的结果如何。然而,除非你的模型是正确的,否则这将产生一个质量极差的数据集。与我的导师Tom Rainforth博士合作,我们将首先致力于加深对这个问题的理解:对导致BED失败的错误规范进行分类;形成诊断故障的指标和避免故障发生的最佳实践;并为故障何时发生提供理论上的保证。在此之后,我们将开发方法来抵消错误规范,理想地将BED的理论优雅和经验性能扩展到我们的模型错误指定的情况下。BED具有巨大的应用潜力,包括量子信息实验、心理学试验、构建逼真的警察草图和指导药物发现。随着这些应用程序变得越来越复杂,模型规范错误变得越来越普遍;因此,进一步调查BED中的错误规范是相关的。该项目属于EPSRC的“统计和应用概率”研究领域。
英文摘要
Much of the research in statistics and machine learning has focussed on methods of analysing data once it has already been acquired, but the question of how to best collect data in the first place has been under-explored. Gathering data can be expensive, therefore practitioners are limited by the amount of data they can collect. When this is done without care, it can produce poor quality data - potentially leading to inaccurate results and incorrect conclusions, regardless of how advanced one's analytical toolkit is. It is therefore vital to endeavour to gather good quality data before analysis. Research in experimental design aims to address this issue, providing practitioners with methods of collecting informative data that will lead to reliable results and strong conclusions. To outline Bayesian experimental design (BED), we consider the following setting: there are several beacons within an area, each emitting a signal, and a practitioner wishes to locate the beacons. The data-gathering process involves the practitioner choosing locations in which to probe the signal, then recording the strength of that signal at these locations. With infinite resources, the practitioner would be able to perfectly locate the beacons, but in practice they are constrained to performing only a finite number of experiments. BED then works as follows: before collecting any data, the practitioner will first form a statistical model of the strength of the signal at a given location in terms of the unknown locations of the beacons, and they will specify their prior beliefs about the locations of the beacons with a prior distribution on these locations. BED procedures can then provide the practitioner with the best locations to probe the signal - where "best" is defined as the locations that will lead to the largest increase in information about the beacon locations. The above can be easily generalised to other settings. BED is both theoretically sound and performs well practically, however, it can break down if our statistical model of the data is misspecified, i.e., if the true data-generating process is different from the model that we specified. Bayesian statistical methodology is always vulnerable to model misspecification, but unfortunately this fact is particularly problematic for experimental design, where we are not just using our model to analyse the data, but also to collect it in the first place. In the worst case, there are models in which the optimal course of action is to pick all your designs in exactly the same place, regardless of the outcomes you observe. However, unless your model is correct, this will produce an extremely poor quality dataset.In collaboration with my supervisor, Dr Tom Rainforth, we will first aim to deepen understanding of this problem: categorising the ways in which misspecification causes BED to fail; forming metrics to diagnose this failure and best practices to avoid its occurrence; and providing theoretical guarantees for when failure will occur. Following this, we will develop methods to counteract misspecification, ideally expanding the theoretical elegance and empirical performance of BED to cases when our model is misspecified. BED has enormous potential application, including quantum information experiments, psychology trials, constructing lifelike police sketches, and guiding drug discovery. As these applications become more complex, model misspecification becomes more prevalent; it is therefore pertinent to further investigate misspecification in BED. This project falls within the EPSRC 'statistics and applied probability' research area.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金