Data Reduction and Large-Scale Inference - Bayesian Coresets
Data Reduction and Large-Scale Inference - Bayesian Coresets
批准号:
2592814
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2021
资助国家:
英国
项目状态:
未结题
起止时间:
2021 至 --
中文摘要
在大规模数据环境中使用贝叶斯方法是有吸引力的,因为它们提供了一致的不确定性量化和先前的规范。不幸的是,贝叶斯推理算法通常在计算上是不可伸缩的,这使得它们在大数据集上的应用变得困难或不可行。随着现代数据集的不断扩大,推理过程必须具有可伸缩性,同时保持对其结果质量的理论保证。于是,问题自然出现了:如何有原则地减少数据,以某种方式从海量的高维数据集中提取有意义的结构,并将其浓缩为更小、更低维的数据集,从而降低分析成本。以前关于扩展贝叶斯推理的工作集中在增强算法,例如,在每次迭代中仅使用随机数据子样本。然而,通过利用数据往往是冗余的这一观点,最近关于贝叶斯核集的工作已经提供了许多方法来寻找比原始数据集小得多的数据的加权子集(称为CoReset)。然后,这种重置可以在许多现有的后验推断算法中不加改变地被利用,从而提供计算加速并保证后验逼近误差。通过确保CoReset构造加上从CoReset估计后续回归参数的组合成本小于从整个数据集估计推断参数的成本,可以实现显著的计算收益。这些想法也可以扩展到其他应用,例如贝叶斯推断,其中不使用点估计,而是使用MCMC或SMC技术从后验分布中采样参数。这样的采样过程涉及重复评估似然函数,该函数使用较小的重置比使用完整数据集的成本更低。在这个主题上的工作也可以朝着将共重置方法与非线性降维技术相结合的方向进行。这种技术的设计不是为了减少数据点的数量,而是通过识别和利用这样一个事实,即数据可能集中在嵌入在高维空间中的低本征维的流形周围,从而减少每个数据点的维度。还有其他几个有趣的研究方向可能会采取这项工作;当前的核心重置减少方法依赖于数据点的完全或条件独立。这些方法可以在多大程度上扩展到这个制度之外?降维方法可以放在一个有充分依据和统一的概率框架内吗?大学的项目主管将是尼克·怀特利和罗伯特·艾利森。“工业”副主管(S)将来自NCSC内的机器学习研究组,该研究组完全致力于大规模贝叶斯推理技术的研究,包括数据简化方法,并将加入我们定期的详细技术讨论。该小组与英国大学研究界在data-science/computational-statistics/machine-learning领域以及与艾伦·图灵研究所和NCSC研究活动有很好的联系
英文摘要
The use of Bayesian methods in large-scale data settings is attractive due to the coherent uncertainty quantification, and prior specification they provide. Unfortunately, Bayesian inference algorithms are not generally computationally scalable, making their application to large datasets difficult or infeasible. As modern data sets continue to grow ever larger, it is essential for inference procedures to be scalable whilst retaining theoretical guarantees on the quality of their results. The question then naturally arises of how to reduce data in a principled manner, somehow extracting the meaningful structure in massive, high-dimensional data sets and condensing it into a smaller, lower-dimensional data sets which are less costly to analyse. Previous work on scaling Bayesian inference has focused on augmenting algorithms to, for example, use only a random data subsample at each iteration. However, by leveraging the insight that data is often redundant, recent work on Bayesian coresets has provided numerous approaches to finding a weighted subset of the data (called a coreset) that is much smaller than the original dataset. This coreset can then be exploited in many existing posterior inference algorithms without alteration, providing computational speedup and guarantees on posterior approximation error. Significant computational gains can be achieved by ensuring that the combined cost of coreset construction plus follow-on regression-parameter estimation from the coreset is less than that of estimating the inference parameters from the full dataset. These ideas can extend to other applications too, for example to Bayesian inference where, rather than using point estimates, parameters are sampled from a posterior distribution using MCMC or SMC techniques. Such sampling processes involve repeatedly evaluating the likelihood function which is less costly using a small coreset than it is for the full dataset. Work on this topic could also be taken in the direction hybridizing coreset methods with nonlinear dimensionality reduction techniques. Such techniques are designed not to reduce the number of data points, but rather the dimension of each data point, by recognizing and exploiting the fact that data may be concentrated around a manifold of low intrinsic dimension, embedded in a high-dimensional space. There are several other interesting research directions in which the work might be taken; current coreset reduction methods rely on full or conditional independence of data points. To what extent can the methods be extended beyond this regime? Can dimensionality reduction methods be placed within a well-founded and unified probabilistic framework? The University project supervisors will be Nick Whiteley and Robert Allison. "Industrial" co-supervisor(s) will be from the machine learning research group within the NCSC which is fully engaged on research into largescale Bayesian inference techniques, including data-reduction methods, and will join with our regular detailed technical discussions. This group is well connected across the UK university research community in the areas of data-science/computational-statistics/machine-learning as well as with the Alan Turing Institute and with NCSC research activities
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
兼捕减少装置(Bycatch Reduction Devices, BRD)对拖网网囊系统水动力及渔获性能的调控机制
-
批准号:32373187
-
项目类别:面上项目
-
资助金额:50万元
-
批准年份:2023
-
负责人:唐浩
-
依托单位: