课题基金 / 基金详情

From trivial representations to learning concepts in AI by exploiting unique data

From trivial representations to learning concepts in AI by exploiting unique data
通过利用独特的数据,从琐碎的表示到学习人工智能中的概念
批准号:
EP/X017680/1
负责人:
Sotirios Tsaftaris
金额:
$25.78万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2023
资助国家:
英国
项目状态:
未结题
起止时间:
2023 至 --

项目摘要

项目成果

Sotirios Tsaftaris的其他基金

相似基金

相关文献

中文摘要
翻译
一场基于人工智能的革命及其社会经济效益的前景是诱人的。我们希望生活在这样一个世界里:人工智能能够高效地学习,表现优异,风险最小。这样的世界是非常令人兴奋的。我们倾向于相信人工智能从数据中学习更高层次的概念,但事实并非如此。特别是在图像等数据中,即使提供了数百万个示例,人工智能也会从数据中提取相当琐碎(低级)的概念。我们经常听说,提供更多高多样性的数据应该有助于提高人工智能可以提取的信息。这种数据积累确实涉及隐私和成本问题。实际上,预处理和净化数据(即删除不需要的信息)的需要也带来了相当大的成本。但更关键的是,在一些关键应用(例如医疗保健)中,某些事件(例如疾病)可能是罕见的或真正独特的。收集越来越多的数据并不会改变这种罕见数据的相对频率。目前的人工智能似乎并不具有数据效率:它很难利用独特和稀有数据中存在的信息金矿。该项目旨在回答一个关键的研究问题:**为什么人工智能与概念斗争,独特数据的作用是什么?**我们怀疑人工智能在概念上挣扎有几个原因:A)我们用来从数据中提取信息的机制(称为表示学习)依赖于非常简单的假设,而这些假设并不能反映真实数据在世界上的存在方式。例如,我们知道数据之间存在相关性,而我们现在做出了根本没有相关性的简化假设。我们建议在我们想要提取的概念中引入更强的因果关系假设。这应该反过来帮助我们提取更好的信息。B)要学习任何模型,我们必须使用优化过程来找到模型的参数。我们在这些过程中发现了一个弱点:独特和罕见的数据不会得到如此多的关注,或者如果它们得到了一些关注,那也是偶然的。这导致在提取信息时出现相当大的不一致性。此外,有时会提取出错误的信息,要么是因为我们发现了次优表示,要么是因为我们锁定了一些从清理过程中逃脱的数据——因为没有这样完美的过程总是可以保证的。我们想要理解为什么存在这种不一致,并建议设计一些方法,以确保在训练模型时,即使从罕见的数据中也能始终如一地提取信息。B和a之间有紧密的联系。如果没有更好地优化学习函数的新方法,我们就不能从稀有数据中可靠地提取表征,因此我们就不能强加我们需要的因果关系。关于这项工作,还有一个额外的因素有助于回答问题的第二部分。罕见和独特的数据实际上可能揭示独特的因果关系。这是一个非常诱人的前景,我们提出的工作旨在调查。我们建议的工作有相当大的和广泛的回报。我们在这里提出了人工智能的基础,因为它是数据高效的,不应该需要盲目地收集数据,因为这给公众带来了所有的隐私担忧。因为它学习了高层次的概念,它将更熟练地授权决策工具,以支持如何达成决策。因为我们在提取这些概念时引入了强大的因果先验,我们减少了学习琐碎数据关联的风险。总的来说,人工智能研究界的一个主要目标是创造出一种人工智能,它可以泛化到训练期间可用的新数据之外的未知数据。我们希望我们的人工智能将使我们更接近这一目标,从而进一步为人工智能在现实世界中的更广泛部署铺平道路。
英文摘要
The prospect of an AI-based revolution and its socio-economic benefits is tantalising. We want to live in a world where AI learns effectively with high performance and minimal risks. Such a world is extremely exciting. We tend to believe that AI learns higher level concepts from data, but this is not what happens. Particularly in data such as images, AI extracts rather trivial (low-level) notions from the data even when provided with millions of examples. We often hear that providing more data with high diversity should help improve the information that AI can extract. This data amassing does have though privacy and cost implications. Indeed, considerable cost comes also by the need to pre-process and to sanitise data (i.e. remove unwanted information). More critically, though, in several key applications (e.g. healthcare) some events (e.g. disease) can be rare or truly unique. Collecting more and more data will not change the relative frequency of such rare data. It appears that current AI is not data efficient: it poorly leverages the goldmine of information present in unique and rare data.This project aims to answer a key research question: **Why does AI struggle with concepts, and what is the role of unique data? **We suspect there are several reasons why AI struggles with concepts: A) The mechanisms we use to extract information from data (known as representation learning) rely on very simple assumptions that do not reflect how real data exist in the world. For example, we know that data have correlations, and we now make simplified assumptions of no correlation at all. We propose to introduce stronger assumptions of causal relationships in the concepts we want to extract. This should in turn help us extract better information. B) To learn any model, we do have to use optimisation processes to find the parameters of the model. We find a weakness in these processes: data that are unique and rare do not get so much attention, or if they do get some, it happens by chance. This leads to considerable inconsistency in the extraction of information. In addition, sometimes wrong information is extracted, either because we found suboptimal representations or because we latched on some data that escaped from the sanitisation process -since no such perfect process can always be guaranteed. We want to understand why such inconsistency exists and propose to devise methods that can ensure that when we train models, we can consistently extract information even from rare data.There is a tight connection between B and A. Without new methods that better optimise learning functions we cannot extract representations reliably from rare data, and hence we cannot impose the causal relationships we need. There is an additional element about this work that helps answer the second part of the question. Rare and unique data may actually reveal unique causal relationships. This is a very tantalising prospect that the work we propose aims to investigate. There are considerable and broad rewards of the work we propose. We put herein the underpinnings for an AI that, because it is data efficient, should not require blind amassing of data with all the privacy fears this engenders for the general public. Because it learns high-lever concepts it will be more adept to empower decision tools that can support how decisions have been reached. And because we introduce strong causal priors in extracting these concepts, we reduce the risk of learning trivial data associations. Overall, a major goal of the AI research community is to create AI that can generalise to new unseen data beyond what was available during training time. We hope that our AI will bring us closer to this goal, thus further paving the way to broader deployment of AI to the real world.
期刊论文(3)
专著(0)
科研奖励(0)
会议论文
Unveiling Fairness Biases in Deep Learning-Based Brain MRI Reconstruction
揭示基于深度学习的脑 MRI 重建中的公平偏差
DOI: 10.48550/arxiv.2309.14392
发表时间: 2023
期刊:
影响因子: --
作者: [Du Y]
通讯作者: Du Y
Deep Generative Models - Third MICCAI Workshop, DGM4MICCAI 2023, Held in Conjunction with MICCAI 2023, Vancouver, BC, Canada, October 8, 2023, Proceedings
深度生成模型 - 第三届 MICCAI 研讨会,DGM4MICCAI 2023,与 MICCAI 2023 同期举行,加拿大不列颠哥伦比亚省温哥华,2023 年 10 月 8 日,会议记录
DOI: 10.1007/978-3-031-53767-7_1
发表时间: 2024
期刊:
影响因子: --
作者: [Fernandez V]
通讯作者: Fernandez V
CHAI - EPSRC AI Hub for Causality in Healthcare AI with Real Data
  • 批准号:
    EP/Y028856/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $1311.0万
  • 财政年份:
    2024
  • 负责人:
    Sotirios Tsaftaris
  • 依托单位:
CardiacA.I.: Machine learning for the analysis of multimodal cardiac MR images used in the diagnosis of coronary heart disease
  • 批准号:
    EP/P022928/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $12.86万
  • 财政年份:
    2017
  • 负责人:
    Sotirios Tsaftaris
  • 依托单位:
海外基金