课题基金 / 基金详情

Learning Unconstrained Human Pose Estimation from Low-cost Approximate Annotation

Learning Unconstrained Human Pose Estimation from Low-cost Approximate Annotation
从低成本近似注释学习无约束人体姿势估计
批准号:
EP/H035885/1
负责人:
Mark Everingham
金额:
$12.81万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2010
资助国家:
英国
项目状态:
已结题
起止时间:
2010 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
这项研究是在计算机视觉制造计算机领域进行的,这种计算机可以理解照片和视频中发生的事情。作为人类,我们对其他人类着迷,并捕捉他们活动的无穷无尽的图像,例如我们全家度假的照片,体育赛事的视频,或者闭路电视拍摄的市中心的人们。一台能够理解人们在这类图像中做什么的计算机将能够为我们做很多工作,例如找到我们的孩子挥手的照片,在足球比赛中快进到球门,或者在街上有人打架时发现。实现这些目标的一个基本任务是让计算机理解一个人的姿势--他们如何站立,手臂是否抬起,他们指向哪里?这个姿势估计问题对人类来说很容易,但对于计算机来说却非常困难,因为人们的姿势、体型和穿着的衣服有很大的不同。许多工作都试图解决这个问题,并且在特别的环境下效果很好,例如,人们穿着带有记号笔的特殊西装来帮助寻找肢体,但不适用于真实世界的照片,因为它使用的是简单的棒人模型。我们将通过向计算机展示许多示例图片来教授计算机,从而研究更好的人类外观模型。这种从图片中学习而不是手动构建模型的方法正在显示出很大的进步,但需要示例图片,其中姿势已由人工注释员标记或注释。因为对图片进行注释既慢又麻烦,目前的方法只能处理几百张图片,而且这还不足以了解人类的所有表现方式。我们将克服这个问题,只用一种非常快的方式粗略地标注图片,这样我们就可以用低成本标注大量的图片。然后,我们将开发一些方法,让计算机可以从这个粗略的注释中学习,通过结合许多我们已经知道的图片和信息,比如人体是如何组合在一起的,计算出相应的准确注释。通过学习大量的图像,以及使用粗略注释的方法,我们将能够制作出更强大的模型,展示人类在改变姿势时的样子。这将导致姿势估计方法在现实世界中更好地发挥作用,并有助于从照片和视频了解人类活动的长期目标。
英文摘要
This research is in the area of computer vision - making computers which can understand what is happening in photographs and video. As humans we are fascinated by other humans, and capture endless images of their activities, for example photographs of our family on holiday, video of sports events or CCTV footage of people in a town center. A computer capable of understanding what people are doing in such images would be able to do many jobs for us, for example finding photos of our child waving, fast forwarding to a goal in a football game, or spotting when someone starts a fight in the street. A fundamental task in achieving such aims is to get the computer to understand a person's pose - how are they standing, is their arm raised, where are they pointing? This pose estimation problem is easy for humans but very difficult for computers because people vary so much in their pose, their body shape and the clothing they wear.Much work has tried to solve this problem, and works well in particular settings for example where people wear a special suit with markers to help find the limbs, but does not work for real-world pictures because it uses simple stick man models of humans. We will investigate better models of how humans look by teaching the computer by showing it many example pictures. This approach of learning from pictures instead of building models by hand is showing great progress, but needs example pictures where the pose has been marked or annotated by a human annotator. Because annotating pictures is slow and tiresome current methods make do with a few hundred pictures and this isn't enough to learn all the ways a human can appear. We will overcome this problem by annotating pictures only roughly in a way which is very fast so we can annotate lots of pictures with low cost. We will then develop methods where the computer can learn from this rough annotation, working out what the corresponding exact annotation would be by combining many pictures and information we already know such as how the human body is put together.By having lots of images to learn from, and methods for making use of rough annotation, we will be able to make stronger models of how humans look as they change their pose. This will lead to pose estimation methods which work better in the real world and contribute to longer-term aims in understanding human activity from photographs and video.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
海外基金