课题基金 / 基金详情

Data Reduction and Large-Scale Inference - Bayesian Coresets

Data Reduction and Large-Scale Inference - Bayesian Coresets
数据缩减和大规模推理 - 贝叶斯核心集
批准号:
2592814
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2021
资助国家:
英国
项目状态:
未结题
起止时间:
2021 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
The use of Bayesian methods in large-scale data settings is attractive due to the coherent uncertainty quantification, and prior specification they provide. Unfortunately, Bayesian inference algorithms are not generally computationally scalable, making their application to large datasets difficult or infeasible. As modern data sets continue to grow ever larger, it is essential for inference procedures to be scalable whilst retaining theoretical guarantees on the quality of their results. The question then naturally arises of how to reduce data in a principled manner, somehow extracting the meaningful structure in massive, high-dimensional data sets and condensing it into a smaller, lower-dimensional data sets which are less costly to analyse. Previous work on scaling Bayesian inference has focused on augmenting algorithms to, for example, use only a random data subsample at each iteration. However, by leveraging the insight that data is often redundant, recent work on Bayesian coresets has provided numerous approaches to finding a weighted subset of the data (called a coreset) that is much smaller than the original dataset. This coreset can then be exploited in many existing posterior inference algorithms without alteration, providing computational speedup and guarantees on posterior approximation error. Significant computational gains can be achieved by ensuring that the combined cost of coreset construction plus follow-on regression-parameter estimation from the coreset is less than that of estimating the inference parameters from the full dataset. These ideas can extend to other applications too, for example to Bayesian inference where, rather than using point estimates, parameters are sampled from a posterior distribution using MCMC or SMC techniques. Such sampling processes involve repeatedly evaluating the likelihood function which is less costly using a small coreset than it is for the full dataset. Work on this topic could also be taken in the direction hybridizing coreset methods with nonlinear dimensionality reduction techniques. Such techniques are designed not to reduce the number of data points, but rather the dimension of each data point, by recognizing and exploiting the fact that data may be concentrated around a manifold of low intrinsic dimension, embedded in a high-dimensional space. There are several other interesting research directions in which the work might be taken; current coreset reduction methods rely on full or conditional independence of data points. To what extent can the methods be extended beyond this regime? Can dimensionality reduction methods be placed within a well-founded and unified probabilistic framework? The University project supervisors will be Nick Whiteley and Robert Allison. "Industrial" co-supervisor(s) will be from the machine learning research group within the NCSC which is fully engaged on research into largescale Bayesian inference techniques, including data-reduction methods, and will join with our regular detailed technical discussions. This group is well connected across the UK university research community in the areas of data-science/computational-statistics/machine-learning as well as with the Alan Turing Institute and with NCSC research activities
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
兼捕减少装置(Bycatch Reduction Devices, BRD)对拖网网囊系统水动力及渔获性能的调控机制
  • 批准号:
    32373187
  • 项目类别:
    面上项目
  • 资助金额:
    50万元
  • 批准年份:
    2023
  • 负责人:
    唐浩
  • 依托单位: