Visual AI: An Open World Interpretable Visual Transformer
Visual AI: An Open World Interpretable Visual Transformer
批准号:
EP/T028572/1
负责人:
Andrew Zisserman
金额:
$753.32万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2020
资助国家:
英国
项目状态:
未结题
起止时间:
2020 至 --
中文摘要
点击翻译按钮获取中文摘要
英文摘要
With the advent of deep learning and the availability of big data, it is now possible to train machine learning algorithms for a multitude of visual tasks, such as tagging personal image collections in the cloud, recognizing faces, and 3D shape scanning with phones. However, each of these tasks currently requires training a neural network on a very large image dataset specifically collected and labelled for that task. The resulting networks are good experts for the target task, but they only understand the 'closed world' experienced during training and can 'say' nothing useful about other content, nor can they be applied to other tasks without retraining, nor do they have an ability to explain their decisions or to recognise their limitations. Furthermore, current visual algorithms are usually 'single modal', they 'close their ears' to the other modalities (audio, text) that may be readily available.The core objective of the Programme is to develop the next generation of audio-visual algorithms that does not have these limitations. We will carry out fundamental research to develop a Visual Transformer capable of visual analysis with the flexibility and interpretability of a human visual system, and aided by the other 'senses' - audio and text. It will be able to continually learn from raw data streams without requiring the traditional 'strong supervision' of a new dataset for each new task, and deliver and distill semantic and geometric information over a multitude of data types (for example, videos with audio, very large scale image and video datasets, and medical images with text records).The Visual Transformer will be a key component of next generation AI, able to address multiple downstream audio-visual tasks, significantly superseding the current limitations of computer vision systems, and enabling new and far reaching applications.A second objective addresses transfer and translation. We seek impact in a variety of other academic disciplines and industry which today greatly under-utilise the power of the latest computer vision ideas. We will target these disciplines to enable them to leapfrog the divide between what they use (or do not use) today which is dominated by manual review and highly interactive analysis frame-by-frame, to a new era where automated visual analytics of very large datasets becomes the norm. In short, our goal is to ensure that the newly developed methods are used by industry and academic researchers in other areas, and turned into products for societal and economic benefit. To this end open source software, datasets, and demonstrators will be disseminated on the project website.The ubiquity of digital images and videos means that every UK citizen may potentially benefit from the Programme research in different ways. One example is smart audio-visual glasses, that can pay attention to a person talking by using their lip movements to mask out other ambient sounds. A second is an app that can answer visual questions (or retrieve matches) for text-queries over large scale audio-visual collections, such as a person's entire personal videos. A third is AI-guided medical screening, that can aid a minimally trained healthcare professional to perform medical scans.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1109/iccv48922.2021.00175
发表时间:
2021-04
期刊:
2021 IEEE/CVF International Conference on Computer Vision (ICCV)
影响因子:
--
作者:
[Max Bain;Arsha Nagrani;Gül Varol;Andrew Zisserman]
通讯作者:
Max Bain;Arsha Nagrani;Gül Varol;Andrew Zisserman
WhisperX: Time-Accurate Speech Transcription of Long-Form Audio
WhisperX:长格式音频的时间精确语音转录
DOI:
10.21437/interspeech.2023-78
发表时间:
2023
期刊:
影响因子:
--
作者:
[Bain M]
通讯作者:
Bain M
DOI:
10.1016/j.media.2022.102630
发表时间:
2022-10-09
期刊:
MEDICAL IMAGE ANALYSIS
影响因子:
10.9
作者:
[Alsharid, Mohammad, Cai, Yifan, Noble, J. Alison]
通讯作者:
Noble, J. Alison
PASS: An ImageNet replacement for self-supervised pretraining without human
PASS:ImageNet 替代自监督预训练,无需人工干预
DOI:
--
发表时间:
2021
期刊:
影响因子:
--
作者:
[Asano, Y]
通讯作者:
Asano, Y
DOI:
10.1109/cvpr52729.2023.01213
发表时间:
2022-11
期刊:
2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
作者:
[Titas Anciukevicius;Zexiang Xu;Matthew Fisher;Paul Henderson;Hakan Bilen;N. Mitra;Paul Guerrero]
通讯作者:
Titas Anciukevicius;Zexiang Xu;Matthew Fisher;Paul Henderson;Hakan Bilen;N. Mitra;Paul Guerrero
共 6 条
Seebibyte: Visual Search for the Era of Big Data
-
批准号:EP/M013774/1
-
项目类别:Research Grant
-
资助金额:$569.27万
-
财政年份:2015
-
负责人:Andrew Zisserman
-
依托单位:
Learning to Recognise Dynamic Visual Content from Broadcast Footage
-
批准号:EP/I012001/1
-
项目类别:Research Grant
-
资助金额:$63.82万
-
财政年份:2011
-
负责人:Andrew Zisserman
-
依托单位:
国内基金
海外基金
登录
查看更多内容
基于协同创新视角下AI赋能课程体系的模块化开发与应用研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:吴惠玲
-
依托单位:
基于AI驱动的教育教学平台系统的开发与应用
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:曹琪敏
-
依托单位:
基于AI智链驱动的跨境电商平台系统开发
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:蔡永林
-
依托单位:
AI赋能未成年人心理健康应用研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:傅绪荣
-
依托单位:
备多分AI智能研学系统开发
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:常直杨
-
依托单位:
AI智慧体育操场的设计与应用
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:王斌
-
依托单位:
长沙软件园 “轻量化AI大模型矩阵 ”科技型企业孵化器建设
-
批准号:2026ZYT011
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:方永强
-
依托单位:
面向AI驱动的信息化工程监管与自动化测试平台研发
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:刘登志
-
依托单位:
建筑-音乐跨模态AI生成平台研发与应用
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:许蕴彰
-
依托单位:
适用于AI眼镜的横向错位光学变焦系统技术开发
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:窦健泰
-
依托单位: