EAGER: Construction of Social Interactions in 3D Space from First-Person Videos
EAGER: Construction of Social Interactions in 3D Space from First-Person Videos
批准号:
1651389
负责人:
Jianbo Shi
金额:
$20.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-09-01 至 2018-08-31
中文摘要
对于现实和复杂的人类社会互动,目前还没有精确的建模工具。第一人称视频提供了一个独特的机会,以前所未有的精确度捕捉社交互动。相比之下,目前的第三人称监控视频只能被动地记录互动的少数距离视图,其空间分辨率大大降低。这个探索性研究项目建议利用多个第一人称摄像机作为一个集体工具来捕捉、建模和预测社会行为。提出的研究改变了我们构建现实社会互动模型的方式,同时也推进了第一人称视频识别。如果成功,设想的计算模型可以充当教练,了解成功的交互和失败的构成,从而能够找到解决方案来调解和防止潜在的冲突。拟议的研究将从多个个人角度模拟三维空间中的动态社会互动。识别和预测复杂的社会群体互动是具有挑战性的,因为群体中的人可能有意或无意地采取意想不到的行动。此外,由于个人的喜好和能力的差异,同样的活动可以以不同的方式进行。第一人称视频可能会高度紧张,导致视野中的物体运动迅速且不可预测。基于PI最近使用第一人称摄像机为社会(人-人)和个人(人-场景)互动建模建立的计算基础,本研究将探索社会注意和角色之间的二元性的新概念:社会注意为识别社会角色提供线索,而社会角色促进动态社会形态变化及其相关社会注意的预测。3D模型的形式基础是建立一个视觉记忆,以三种形式存储第一人称社会经验:(a)几何社会形态,(b)第一人称视角的视觉图像,(c)附近第三人称视角看到的第一人称视角。作为概念验证,捕捉社会互动的3D空间模型将在协作社会任务中进行测试,例如组装(宜家)家具,或与一群朋友一起建造积木屋。本研究将构建一个捕获社会互动的标记数据集,并对社会角色识别的准确性和预测社会互动中成员空间运动的精度进行分析。该项目的结果,包括论文和数据集,将通过我们的项目网站(http://www.cis.upenn.edu/~jshi/NSF_SocialMemory/nsf_social_visual_memory.html)向公众发布。在这个项目下创建的软件将通过GitHub向公众提供,GitHub是一个基于web的Git存储库托管服务
英文摘要
Precision modeling tools for realistic and complex human social interaction are not available today. First-person videos provide a unique opportunity to capture social interaction at unprecedented precision. In contrast, current third person surveillance video only records the few distance views of the interaction passively at a much reduced spatial resolution. This exploratory research project proposes to harness multiple first-person cameras as one collective instrument to capture, model, and predict social behaviors. The proposed research transforms the way we construct realistic social interaction models, while also advancing first-person video recognition. If successful, the envisioned computational model can act as a coach who learns what constitutes successful interactions and failures, thus being able to find solutions to mediate and prevent potential conflicts. The proposed research will model dynamic social interactions in 3D space from multiple personal perspectives. Recognition and prediction of complex social group interactions are challenging because people in the group can carry out unexpected actions intentionally or by mistake. In addition, due to variances in individuals' preferences and abilities, the same activities could be carried out in different ways. First-person videos can be highly jittery, resulting in fast and unpredictable object motions in the field of view. Building on PI's recent work establishing computational foundations for modeling social (people-people) and personal (people-scene) interactions using first-person cameras, this research will explore the novel concept the duality between social attention and roles: social attention provides a cue for recognizing social roles, and social roles facilitate the predictions of dynamic social formation change and its associated social attention. The formal foundation of the 3D model is based on constructing a visual memory that stores first-person social experiences in three forms: (a) geometric social formation, (b) visual image of first-person view, and (c) first-person seen by nearby third person views. As a proof-of-concept, the 3D space model capturing social interactions will be tested on collaborative social tasks such as assembling (Ikea) furniture, or building a block house with a group of friends. This research will construct a labeled dataset capturing the interactions, and perform analysis on both accuracy in recognizing social roles and precision in predicting spatial movements of the members in that social interaction. The results of this project, including papers and dataset, will be disseminated to the public through our project website (http://www.cis.upenn.edu/~jshi/NSF_SocialMemory/nsf_social_visual_memory.html). The software created under this project will be made available to the public through GitHub, a web-based Git repository hosting service
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: 1st Sino-USA Summer School in Vision, Learning, Pattern Recognition, VLPR 2009
-
批准号:0940840
-
项目类别:Standard Grant
-
资助金额:$2.45万
-
财政年份:2009
-
负责人:Jianbo Shi
-
依托单位:
RI-Medium: From Actors To Actions: Analysis And Alignment Of Images, Video And Text
-
批准号:0803538
-
项目类别:Standard Grant
-
资助金额:$0.0万
-
财政年份:2008
-
负责人:Jianbo Shi
-
依托单位:
CAREER: Learning to See - A Unified Segmentation and Recognition Approach
-
批准号:0447953
-
项目类别:Standard Grant
-
资助金额:$40.0万
-
财政年份:2005
-
负责人:Jianbo Shi
-
依托单位:
RR:MACNet: Mobile Ad-hoc Camera Networks
-
批准号:0423891
-
项目类别:Continuing Grant
-
资助金额:$19.86万
-
财政年份:2004
-
负责人:Jianbo Shi
-
依托单位:
国内基金
海外基金
Data-driven Recommendation System Construction of an Online Medical Platform Based on the Fusion of Information
-
批准号:--
-
项目类别:外国青年学者研究基金项目
-
资助金额:--
-
批准年份:2024
-
负责人:江洋子
-
依托单位: