课题基金 / 基金详情

New challenges in robust statistical learning

New challenges in robust statistical learning
稳健统计学习的新挑战
批准号:
EP/V002694/1
负责人:
Timothy Cannings
金额:
$33.94万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2021
资助国家:
英国
项目状态:
未结题
起止时间:
2021 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
近年来,我们收集、存储和处理海量数据的能力,加上技术的快速进步,导致了数据驱动决策的广泛采用。这包括新的应用领域,如精准医学,医生正在使用数据来告知他们的诊断和治疗建议。在金融等其他领域,银行使用大量历史数据来决定新客户是否有可能(或不会)拖欠贷款。通常的情况是,我们需要根据与现有患者相关的一些(训练)数据,对一些未来的患者或客户做出离散的预测。在统计学中,这类问题被称为分类问题。许多分类方法都是建立在假设我们未来可能遇到的任何数据与我们的训练数据具有相同的分布的基础上的。当然,这一假设并不总是有效的--与一组患者或客户相关的数据不一定遵循与来自一组新人的数据相同的分布。在这项研究中,我们将开发新的稳健的分类算法,能够处理噪声和不完整的数据。特别是,新的方法将使从业者能够合并多个噪声数据来源,对现有方法提出修改建议,以确保它们对数据中的损坏具有健壮性,并引入新的方法来克服丢失数据造成的问题。我们还将从理论上对决策算法在面对噪声、损坏和不完整数据时的局限性提供新的理论理解。我们的新方法将在许多情况下适用:-我们可能从特定地点(实验室或医院)的患者那里收集数据,但希望在不同的地点进行预测。-我们可能无法访问完整的数据集。例如,出于隐私原因,用户可能不会泄露其某些个人信息。在其他设置中,我们可能需要通过删除一些识别协变量来匿名数据。-通常情况下,涉及的数据类型的复杂性将意味着我们观察不到真实的数据。相反,我们只能访问数据的近似值。这通常发生在现代环境中,从业者使用亚马逊机械土耳其人等众包服务来标记他们的数据--这样的服务很少是完全准确的。-对手可能能够任意污染一小部分数据(例如,通过在线执行人工活动)。我们的工作将使从业者能够利用目前不适合使用的数据。我们还将提供对特定目的最有用的数据类型的新见解。
英文摘要
In recent years, our ability to collect, store and process vast amounts of data, coupled with rapid advances in technology, have led to the widespread adoption of data-driven decision-making. This includes new application areas, such as precision medicine, where doctors are using data to inform their diagnoses and treatment recommendations. In other areas, such as finance, banks use huge amounts of historical data in order to decide whether a new customer is likely (or not) to default on their loan repayments. It is often the case that we are required to make a discrete prediction about some future patient or customer, based on some (training) data relating to existing patients. In statistics, problems of this type are called classification problems. Many methods for classification are built on the assumption that any future data we may encounter has the same distribution as our training data. Of course, this assumption is not always valid -- data relating to one set of patients or customers will not necessarily follow the same distribution as data from a new set of people. In this research, we will develop new robust classification algorithms that can deal with noisy and incomplete data. In particular, the new methodology will enable practitioners to combine multiple sources of noisy data, propose modifications to existing methods in order to guarantee they are robust to corruptions in the data, and introduce novel ways of overcoming the issues caused by missing data. We will also provide new theoretical understanding of the limitations of decision-making algorithms when faced with noisy, corrupted and incomplete data. There are a number of scenarios where our new approaches will be applicable: - We may have data collected from patients in a particular location (lab or hospital) but wish to make predictions in a different location.- We may not have access to the full dataset. For example, for privacy reasons, uses may not disclose some of their personal information. In other settings, we may be required to anonymise the data by removing some identifying covariates. - Often the complexity of the type of data involved will mean that we don't observe the true data. Instead, we only have access to an approximation of the data. This typically occurs in modern settings, where practitioners use crowd-sourcing services such as the Amazon Mechanical Turk to label their data -- such services are rarely perfectly accurate. - It may be that an adversary is able to arbitrarily contaminate a small proportion of the data (for instance by performing artificial activity online).Our work will enable practitioners to utilise data that is currently not appropriate for use. We will also provide new insight into the kinds of data that are most useful for a particular purpose.
期刊论文(9)
专著(0)
科研奖励(0)
会议论文
Optimal subgroup selection
最优子组选择
DOI: 10.1214/23-aos2328
发表时间: 2023
期刊: The Annals of Statistics
影响因子: --
作者: [Reeve H]
通讯作者: Reeve H
Trace-class Gaussian priors for Bayesian learning of neural networks with MCMC
使用 MCMC 进行神经网络贝叶斯学习的迹级高斯先验
DOI: 10.1093/jrsssb/qkac005
发表时间: 2023
期刊: Statistical Methodology
影响因子: --
作者: [Sell T]
通讯作者: Sell T
Adaptive transfer learning
自适应迁移学习
DOI: 10.1214/21-aos2102
发表时间: 2021
期刊: The Annals of Statistics
影响因子: --
作者: [Reeve H]
通讯作者: Reeve H
The correlation-assisted missing data estimator
相关辅助缺失数据估计器
DOI: --
发表时间: 2022
期刊: Journal of Machine Learning Research
影响因子: 6
作者: [Cannings T.I.]
通讯作者: Cannings T.I.
国内基金
海外基金
Supply Chain Collaboration in addressing Grand Challenges: Socio-Technical Perspective
  • 批准号:
    --
  • 项目类别:
    外国青年学者研究基金项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    Lim Jia Jia
  • 依托单位:
Navigating Sustainability: Understanding Environm ent,Social and Governanc e Challenges and Solution s for Chinese Enterprises in Pakistan's CPEC Framew ork
  • 批准号:
    --
  • 项目类别:
    外国学者研究基金项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    Noshaba Aziz
  • 依托单位: