Visual AI: An Open World Interpretable Visual Transformer
Visual AI: An Open World Interpretable Visual Transformer
批准号:
EP/T028572/1
负责人:
Andrew Zisserman
金额:
$753.32万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2020
资助国家:
英国
项目状态:
未结题
起止时间:
2020 至 --
中文摘要
随着深度学习的出现和大数据的可用性,现在可以为多种视觉任务训练机器学习算法,例如在云中标记个人图像集合,识别人脸和使用手机进行3D形状扫描。然而,这些任务中的每一个目前都需要在专门为该任务收集和标记的非常大的图像数据集上训练神经网络。由此产生的网络是目标任务的优秀专家,但它们只理解在训练期间经历的“封闭世界”,不能“说”任何有用的内容,也不能在未经重新训练的情况下应用于其他任务,也没有能力解释他们的决定或认识到他们的局限性。此外,目前的视觉算法通常是“单一模式”,他们“关闭耳朵”,以其他形式(音频,文本),可能是现成的。该计划的核心目标是开发下一代的视听算法,没有这些限制。我们将进行基础研究,开发一种能够进行视觉分析的视觉Transformer,具有人类视觉系统的灵活性和可解释性,并得到其他“感官”-音频和文本的帮助。它将能够不断地从原始数据流中学习,而不需要对每个新任务的新数据集进行传统的“强监管”,并在多种数据类型上提供和提取语义和几何信息。(例如,带有音频的视频、超大规模图像和视频数据集以及带有文本记录的医学图像)。视觉Transformer将成为下一代AI的关键组件,能够解决多个下游视听任务,显著取代当前计算机视觉系统的局限性,并实现新的和深远的应用。第二个目标解决传输和翻译。我们寻求在各种其他学科和行业的影响,今天大大利用最新的计算机视觉思想的力量。我们将针对这些学科,使他们能够跨越他们今天使用(或不使用)的内容之间的鸿沟,这些内容由手动审查和高度交互式的逐帧分析主导,进入一个新时代,对非常大的数据集进行自动化可视化分析成为常态。简而言之,我们的目标是确保新开发的方法被其他领域的工业和学术研究人员使用,并转化为具有社会和经济效益的产品。为此,开放源码软件,数据集和演示将在项目网站上传播。数字图像和视频的无处不在意味着每个英国公民都可能以不同的方式从该计划的研究中受益。一个例子是智能视听眼镜,它可以通过使用嘴唇运动来掩盖其他环境声音来关注一个人的谈话。第二个是一个应用程序,可以回答视觉问题(或检索匹配)的文本查询大规模的视听收藏,如一个人的整个个人视频。第三个是人工智能引导的医疗筛查,可以帮助受过最低限度培训的医疗保健专业人员进行医疗扫描。
英文摘要
With the advent of deep learning and the availability of big data, it is now possible to train machine learning algorithms for a multitude of visual tasks, such as tagging personal image collections in the cloud, recognizing faces, and 3D shape scanning with phones. However, each of these tasks currently requires training a neural network on a very large image dataset specifically collected and labelled for that task. The resulting networks are good experts for the target task, but they only understand the 'closed world' experienced during training and can 'say' nothing useful about other content, nor can they be applied to other tasks without retraining, nor do they have an ability to explain their decisions or to recognise their limitations. Furthermore, current visual algorithms are usually 'single modal', they 'close their ears' to the other modalities (audio, text) that may be readily available.The core objective of the Programme is to develop the next generation of audio-visual algorithms that does not have these limitations. We will carry out fundamental research to develop a Visual Transformer capable of visual analysis with the flexibility and interpretability of a human visual system, and aided by the other 'senses' - audio and text. It will be able to continually learn from raw data streams without requiring the traditional 'strong supervision' of a new dataset for each new task, and deliver and distill semantic and geometric information over a multitude of data types (for example, videos with audio, very large scale image and video datasets, and medical images with text records).The Visual Transformer will be a key component of next generation AI, able to address multiple downstream audio-visual tasks, significantly superseding the current limitations of computer vision systems, and enabling new and far reaching applications.A second objective addresses transfer and translation. We seek impact in a variety of other academic disciplines and industry which today greatly under-utilise the power of the latest computer vision ideas. We will target these disciplines to enable them to leapfrog the divide between what they use (or do not use) today which is dominated by manual review and highly interactive analysis frame-by-frame, to a new era where automated visual analytics of very large datasets becomes the norm. In short, our goal is to ensure that the newly developed methods are used by industry and academic researchers in other areas, and turned into products for societal and economic benefit. To this end open source software, datasets, and demonstrators will be disseminated on the project website.The ubiquity of digital images and videos means that every UK citizen may potentially benefit from the Programme research in different ways. One example is smart audio-visual glasses, that can pay attention to a person talking by using their lip movements to mask out other ambient sounds. A second is an app that can answer visual questions (or retrieve matches) for text-queries over large scale audio-visual collections, such as a person's entire personal videos. A third is AI-guided medical screening, that can aid a minimally trained healthcare professional to perform medical scans.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1109/iccv48922.2021.00175
发表时间:
2021-04
期刊:
2021 IEEE/CVF International Conference on Computer Vision (ICCV)
影响因子:
--
作者:
[Max Bain;Arsha Nagrani;Gül Varol;Andrew Zisserman]
通讯作者:
Max Bain;Arsha Nagrani;Gül Varol;Andrew Zisserman
WhisperX: Time-Accurate Speech Transcription of Long-Form Audio
WhisperX:长格式音频的时间精确语音转录
DOI:
10.21437/interspeech.2023-78
发表时间:
2023
期刊:
影响因子:
--
作者:
[Bain M]
通讯作者:
Bain M
DOI:
10.1016/j.media.2022.102630
发表时间:
2022-10-09
期刊:
MEDICAL IMAGE ANALYSIS
影响因子:
10.9
作者:
[Alsharid, Mohammad, Cai, Yifan, Noble, J. Alison]
通讯作者:
Noble, J. Alison
PASS: An ImageNet replacement for self-supervised pretraining without human
PASS:ImageNet 替代自监督预训练,无需人工干预
DOI:
--
发表时间:
2021
期刊:
影响因子:
--
作者:
[Asano, Y]
通讯作者:
Asano, Y
DOI:
10.1109/cvpr52729.2023.01213
发表时间:
2022-11
期刊:
2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
作者:
[Titas Anciukevicius;Zexiang Xu;Matthew Fisher;Paul Henderson;Hakan Bilen;N. Mitra;Paul Guerrero]
通讯作者:
Titas Anciukevicius;Zexiang Xu;Matthew Fisher;Paul Henderson;Hakan Bilen;N. Mitra;Paul Guerrero
共 6 条
Seebibyte: Visual Search for the Era of Big Data
-
批准号:EP/M013774/1
-
项目类别:Research Grant
-
资助金额:$569.27万
-
财政年份:2015
-
负责人:Andrew Zisserman
-
依托单位:
Learning to Recognise Dynamic Visual Content from Broadcast Footage
-
批准号:EP/I012001/1
-
项目类别:Research Grant
-
资助金额:$63.82万
-
财政年份:2011
-
负责人:Andrew Zisserman
-
依托单位:
国内基金
海外基金
登录
查看更多内容
面向AI驱动的信息化工程监管与自动化测试平台研发
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:刘登志
-
依托单位:
建筑-音乐跨模态AI生成平台研发与应用
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:许蕴彰
-
依托单位:
适用于AI眼镜的横向错位光学变焦系统技术开发
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:窦健泰
-
依托单位:
AI赋能中国传统壁画大模型开发与数字再生展示
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:朱亮亮
-
依托单位:
基于协同创新视角下AI赋能课程体系的模块化开发与应用研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:吴惠玲
-
依托单位:
AI赋能未成年人心理健康应用研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:傅绪荣
-
依托单位:
备多分AI智能研学系统开发
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:常直杨
-
依托单位:
带阻尼的弹簧型减振系统的虚拟建模、能控性分析及AI数智教育技术的开发
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:王成强
-
依托单位:
基于大数据分析与AI算力的民营教培企业提档升级内控管理系统研发
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:卞禹臣
-
依托单位:
智能吊篮AI检测盒子开发
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:田申
-
依托单位: