Cross-modal egocentric activity recognition and zero-shot learning
Cross-modal egocentric activity recognition and zero-shot learning
批准号:
1971464
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2017
资助国家:
英国
项目状态:
已结题
起止时间:
2017 至 --
中文摘要
点击翻译按钮获取中文摘要
英文摘要
The availability of low-cost wearable cameras has renewed the interest of first-person human activity analysis. The recognition of first-person activities has important challenges to be addressed such as, rapid changes in illuminations, significant camera motion and complex hand-object manipulations. In recent years, the advances in deep learning have influenced significantly the computer vision community, as convolutional networks gave impressive results in tasks such as, object recognition and detection, scene understanding and image segmentation. Convolutional networks have been used with success in first-person activity recognition as well. Before the emergence of deep learning the community of first-person computer vision was focused on the engineering of important egocentric features that capture properties of the first-person point of view, such as hand-object interactions and gaze. Convolutional networks allow the learning of such features automatically using big amounts of data, eliminating the need of hand-designed features. In this work, we focus on activity recognition with convolutional networks. Influenced by the recent success of multi-stream architectures, we are investigating their applicability in egocentric videos, by employing multiple modalities for training the models. An important observation that motivates us is that humans combine their senses to understand concepts of the world, such as acoustic and visual information. To this end, we propose the employment of both videos and sounds towards more accurate activity recognition. Specifically, we will investigate how shared aligned representations can be learnt using the multi-stream paradigm. Moreover, we are interested in temporal feature pooling methods to leverage information that spans over the whole video, as in many cases the whole video should be observed in order to be able to discriminate between similar actions. Our final goal is to employ these ideas in zero-shot learning. Zero-shot learning is being able to solve a task despite not having received any training examples of that task. An example is to recognize activities without having seen any video of these activities during training. This can be done by using the knowledge of trained classifiers (trained in other classes and not in the ones to be predicted by the zero-shot paradigm) and additional knowledge about the new classes.
期刊论文(4)
专著(0)
科研奖励(0)
会议论文
DOI:
10.1109/icassp39728.2021.9413376
发表时间:
2021-03
期刊:
ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
作者:
[E. Kazakos;Arsha Nagrani;Andrew Zisserman;D. Damen]
通讯作者:
E. Kazakos;Arsha Nagrani;Andrew Zisserman;D. Damen
DOI:
10.1109/tpami.2020.2991965
发表时间:
2021-11-01
期刊:
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
影响因子:
23.6
作者:
[Damen, Dima, Doughty, Hazel, Wray, Michael]
通讯作者:
Wray, Michael
国内基金
海外基金
基于异构医学影像数据的深度挖掘技术及中枢神经系统重大疾病的精准预测
-
批准号:61672236
-
项目类别:面上项目
-
资助金额:64.0万元
-
批准年份:2016
-
负责人:王骏
-
依托单位: