UMPIRE: United Model for the Perception of Interactions in visuoauditory REcognition
UMPIRE: United Model for the Perception of Interactions in visuoauditory REcognition
批准号:
EP/T004991/1
负责人:
Dima Damen
金额:
$127.65万
依托单位:
依托单位国家:
英国
项目类别:
Fellowship
财政年份:
2020
资助国家:
英国
项目状态:
未结题
起止时间:
2020 至 --
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Humans interact with tens of objects daily, at home (e.g. cooking/cleaning) or outdoors (e.g. ticket machines/shopping bags), during working (e.g. assembly/machinery) or leisure hours (e.g. playing/sports), individually or collaboratively. When observing people interacting with objects, our vision assisted by the sense of hearing is the main tool to perceive these interactions. Let's take the example of boiling water from a kettle. We observe the actor press a button, wait and hear the water boil and the kettle's light go off before water is used for, say, preparing tea. The perception process is formed from understanding intentional interactions (called ideomotor actions) as well as reactive actions to dynamic stimuli in the environment (referred to as sensormotor actions). As observers, we understand and can ultimately replicate such interactions using our sensory input, along with our underlying complex cognitive processes of event perception. Evidence in behavioural sciences demonstrates that these human cognitive processes are highly modularised, and these modules collaborate to achieve our outstanding human-level perception.However, current approaches in artificial intelligence are lacking in their modularity and accordingly their capabilities. To achieve human-level perception of object interactions, including online perception when the interaction results in mistakes (e.g. water is spilled) or risks (e.g. boiling water is spilled), this fellowship focuses on informing computer vision and machine learning models, including deep learning architectures, from well-studied cognitive behavioural frameworks.Deep learning architectures have achieved superior performance, compared to their hand-crafted predecessors, on video-level classification, however their performance on fine-grained understanding within the video remains modest. Current models are easily fooled by similar motions or incomplete actions, as shown by recent research. This fellowship focuses on empowering these models through modularisation, a principle proven since the 50s in Fodor's Modularity of the Mind, and frequently studied by cognitive psychologists in controlled lab environments. Modularity of high-level perception, along with the power of deep learning architectures, will bring a new understanding to videos analysis previously unexplored.The targeted perception, of daily and rare object interactions, will lay the foundations for applications including assistive technologies using wearable computing, and robot imitation learning. We will work closely with three industrial partners to pave potential knowledge transfer paths to applications.Additionally, the fellowship will actively engage international researchers through workshops, benchmarks and public challenges on large datasets, to encourage other researchers to address problems related to fine-grained perception in video understanding.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Rescaling Egocentric Vision
重新调整以自我为中心的愿景
DOI:
10.48550/arxiv.2006.13256
发表时间:
2020
期刊:
影响因子:
--
作者:
[Damen D]
通讯作者:
Damen D
DOI:
10.1109/cvpr42600.2020.00095
发表时间:
2019-12
期刊:
2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
作者:
[Hazel Doughty;I. Laptev;W. Mayol-Cuevas;D. Damen]
通讯作者:
Hazel Doughty;I. Laptev;W. Mayol-Cuevas;D. Damen
Computer Vision - ACCV 2022 - 16th Asian Conference on Computer Vision, Macao, China, December 4-8, 2022, Proceedings, Part IV
计算机视觉 - ACCV 2022 - 第十六届亚洲计算机视觉会议,中国澳门,2022 年 12 月 4-8 日,会议记录,第四部分
DOI:
10.1007/978-3-031-26316-3_27
发表时间:
2023
期刊:
影响因子:
--
作者:
[Fragomeni A]
通讯作者:
Fragomeni A
DOI:
10.48550/arxiv.2310.17395
发表时间:
2023-10
期刊:
ArXiv
影响因子:
--
作者:
[Kevin Flanagan;D. Damen;Michael Wray]
通讯作者:
Kevin Flanagan;D. Damen;Michael Wray
Epic-Sounds: A Large-scale Dataset of Actions That Sound
史诗般的声音:声音动作的大规模数据集
DOI:
10.48550/arxiv.2302.00646
发表时间:
2023
期刊:
影响因子:
--
作者:
[Huh J]
通讯作者:
Huh J
共 9 条
LOCATE: LOcation adaptive Constrained Activity recognition using Transfer learning
-
批准号:EP/N033779/1
-
项目类别:Research Grant
-
资助金额:$12.5万
-
财政年份:2016
-
负责人:Dima Damen
-
依托单位:
海外基金