Learning to Recognise Dynamic Visual Content from Broadcast Footage
Learning to Recognise Dynamic Visual Content from Broadcast Footage
批准号:
EP/I011811/1
负责人:
Richard Bowden
金额:
$62.41万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2011
资助国家:
英国
项目状态:
已结题
起止时间:
2011 至 --
中文摘要
这项研究是在计算机视觉领域-使计算机能够理解照片和视频中发生的事情。作为人类,我们对其他人着迷,并捕捉他们活动的无尽图像,例如我们家庭度假的家庭电影,体育赛事的视频或市中心人们的闭路电视镜头。一台能够理解人们在这些图像中所做的事情的计算机将能够为我们做很多工作,例如找到我们孩子挥手的片段,快速前进到足球比赛的目标,或者发现有人在街上打架。对于聋人来说,他们使用的语言结合了手势、面部表情和肢体语言,一台能够在视觉上理解他们行为的计算机将使他们能够用母语进行交流。虽然人类非常善于理解人们在做什么(并且可以学习理解手语等特殊动作),这对计算机来说已经证明是极其具有挑战性的。许多工作都试图解决这个问题,并且在特定的设置中效果很好例如计算机可以判断一个人是否在走路,只要他们清楚地做并且面向侧面,或能听懂几个手语手势,只要配合,慢慢地打。我们将通过向计算机展示许多示例视频来学习更好的识别活动的模型。为了确保我们的方法适用于各种设置,我们将使用电影和电视中的真实的世界视频。对于每个视频,我们必须告诉计算机它代表什么,例如扔球或男人拥抱女人。以这种方式收集和标记大量视频的成本很高,因此我们将从电视可用的字幕文本和脚本中自动提取近似标签。我们的新方法将联合收割机从大量近似标记的视频中学习(便宜是因为我们自动获得标签),使用上下文信息,如人们同时做哪些动作,或者一个动作如何导致另一个动作(他打了那个人,那个人福尔斯倒在地板上),以及用于理解人的姿势的计算机视觉方法(他们是如何站立的),他们是如何移动的,以及他们使用的物体。通过有大量的视频可以学习,以及使用近似标签的方法,我们将能够制作更强大,更灵活的人类活动模型。这将导致识别方法在真实的世界中工作得更好,并有助于解释手语和自动标记视频内容等应用。
英文摘要
This research is in the area of computer vision - making computers which can understand what is happening in photographs and video. As humans we are fascinated by other humans, and capture endless images of their activities, for example home movies of our family on holiday, video of sports events or CCTV footage of people in a town center. A computer capable of understanding what people are doing in such images would be able to do many jobs for us, for example finding clips of our children waving, fast forwarding to a goal in a football game, or spotting when someone starts a fight in the street. For Deaf people, who use a language combining hand gestures with facial expression and body language, a computer which could visually understand their actions would allow them to communicate in their native language. While humans are very good at understanding what people are doing (and can learn to understand special actions such as sign language), this has proved extremely challenging for computers.Much work has tried to solve this problem, and works well in particular settings for example the computer can tell if a person is walking so long as they do it clearly and face to the side, or can understand a few sign language gestures as long as the signer cooperates and signs slowly. We will investigate better models for recognising activities by teaching the computer by showing it many example videos. To make sure our method works well for all kinds of setting we will use real world video from movies and TV. For each video we have to tell the computer what it represents, for example throwing a ball or a man hugging a woman . It would be expensive to collect and label lots of videos in this way, so instead we will extract approximate labels automatically from subtitle text and scripts which are available for TV. Our new methods will combine learning from lots of approximately labelled video (cheap because we get the labels automatically), use of contextual information such as which actions people do at the same time, or how one action leads to another ( he hits the man, who falls to the floor ), and computer vision methods for understanding the pose of a person (how they are standing), how they are moving, and the objects which they are using.By having lots of video to learn from, and methods for making use of approximate labels, we will be able to make stronger and more flexible models of human activities. This will lead to recognition methods which work better in the real world and contribute to applications such as interpreting sign language and automatically tagging video with its content.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Geometric Mining: Scaling Geometric Hashing to Large Datasets
几何挖掘:将几何哈希扩展到大型数据集
DOI:
10.1109/iccvw.2015.135
发表时间:
2015
期刊:
影响因子:
--
作者:
[Gilbert A]
通讯作者:
Gilbert A
DOI:
10.1109/cvpr.2012.6247928
发表时间:
2012
期刊:
影响因子:
--
作者:
[Eng-Jon Ong]
通讯作者:
Eng-Jon Ong
DOI:
10.5244/c.28.108
发表时间:
2014-09
期刊:
影响因子:
--
作者:
[Simon Hadfield;R. Bowden]
通讯作者:
Simon Hadfield;R. Bowden
DOI:
10.1016/j.cviu.2017.02.001
发表时间:
2016-09
期刊:
Comput. Vis. Image Underst.
影响因子:
--
作者:
[Andrew Gilbert;R. Bowden]
通讯作者:
Andrew Gilbert;R. Bowden
Scene particles: unregularized particle-based scene flow estimation.
场景粒子:基于非正则粒子的场景流估计。
DOI:
10.1109/tpami.2013.162
发表时间:
2014
期刊:
IEEE transactions on pattern analysis and machine intelligence
影响因子:
23.6
作者:
[Hadfield S]
通讯作者:
Hadfield S
共 6 条
ROSSINI: Reconstructing 3D structure from single images: a perceptual reconstruction approach
-
批准号:EP/S016317/1
-
项目类别:Research Grant
-
资助金额:$55.84万
-
财政年份:2019
-
负责人:Richard Bowden
-
依托单位:
ExTOL: End to End Translation of British Sign Language
-
批准号:EP/R03298X/1
-
项目类别:Research Grant
-
资助金额:$123.84万
-
财政年份:2018
-
负责人:Richard Bowden
-
依托单位:
LILiR2 - Language Independent Lip Reading
-
批准号:EP/E027946/1
-
项目类别:Research Grant
-
资助金额:$44.67万
-
财政年份:2007
-
负责人:Richard Bowden
-
依托单位:
LTER Cross-site: Collaborative Research: DIRT: A Cross-continental, Experimental Study of Forest Soil Organic Matter and Nitrogen Dynamics
-
批准号:0087010
-
项目类别:Standard Grant
-
资助金额:$6.6万
-
财政年份:2000
-
负责人:Richard Bowden
-
依托单位:
Improvement of Soil and Ecosystem Analysis Laboratory
-
批准号:9151189
-
项目类别:Standard Grant
-
资助金额:$4.11万
-
财政年份:1991
-
负责人:Richard Bowden
-
依托单位:
海外基金