RI-Medium: From Actors To Actions: Analysis And Alignment Of Images, Video And Text
RI-Medium: From Actors To Actions: Analysis And Alignment Of Images, Video And Text
批准号:
0803538
负责人:
Jianbo Shi
金额:
$0.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2008
资助国家:
美国
项目状态:
已结题
起止时间:
2008-09-01 至 2011-08-31
中文摘要
视频剪辑和相应的叙述一起提供了比单独提供更丰富的信息,然而大多数当前的识别系统分别处理视觉和文本信息。 PI专注于学习如何准确和鲁棒地识别视频和文本叙事中的相应动作的任务。特别是,他们专注于人类行为的语义描述。这项研究将对数字图书馆中的视频检索、人类行为建模和视频监控等应用产生广泛的影响。PI的研究将通过强大的自动学习对应关系将计算机视觉、自然语言处理和机器学习中的方法紧密结合起来。通过一组松散对齐的视频-文本注释对(例如电影或电视节目及其相关剧本),任务是学习如何将文本中的动作描述与视频中的动作、对象和演员相关联。这种对应对于使用视觉动作外观的文本的语义基础是必不可少的。最根本的挑战是弥合图像和文本的语义鸿沟:图像描述图像区域的几何关系和属性,而自然语言则以语法结构编码抽象的语义关系。在动作理解的背景下弥合这种语义鸿沟是我们研究工作的重点。最终目标是能够识别视频中的动作并为视频中的动作创建文本描述。虽然这一目标对计算机视觉和自然语言处理都提出了挑战,但它也为这两个研究领域之间开辟了一个令人兴奋的新的、非常富有成效的合作,在这两个领域中,识别任务是通过同时学习和推理来实现的。有关该项目的信息,包括论文、结果、数据库和开源代码,将在http://www.seas.upenn.edu/~jshi/#research上提供。
英文摘要
Video clips and corresponding narrations together provide much richer information than either in isolation, yet most current recognition systems process visual and textual information separately. The PIs focus on the task of learning how to recognize corresponding actions in videos and textual narrative accurately and robustly. In particular, they focus on semantic descriptions of human actions. This research will have broad impact on applications including video retrieval in digital libraries, human behavior modeling, and video surveillance.The PIs' research will tightly couple methods in computer vision, natural-language processing, and machine learning through robust, automatically learned correspondences. With a collection of loosely aligned video-text annotation pairs (such as movies or TV shows with their associated screenplays), the task is to learn how to associate action descriptions in text with actions, objects and actors in videos. This correspondence is essential for semantic grounding of text using visual action appearance. The fundamental challenge is bridging the semantic gap of images and of text: images depict geometrical relationships and properties of image regions, while natural language encodes abstract semantic relationships in grammatical structures. Bridging this semantic gap in the context of action understanding is the focus of our research effort.The eventual goal is to be able to recognize actions in videos and create text description for actions in videos. While this goal challenges both computer vision and natural language processing, it also opens up an exciting new and very fruitful collaboration between the two research areas where the task of recognition is achieved by simultaneous learning and inference in both domains.Information on this project, including papers, results, database and open source codes, will be available at http://www.seas.upenn.edu/~jshi/#research
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
EAGER: Construction of Social Interactions in 3D Space from First-Person Videos
-
批准号:1651389
-
项目类别:Standard Grant
-
资助金额:$20.0万
-
财政年份:2016
-
负责人:Jianbo Shi
-
依托单位:
Collaborative Research: 1st Sino-USA Summer School in Vision, Learning, Pattern Recognition, VLPR 2009
-
批准号:0940840
-
项目类别:Standard Grant
-
资助金额:$2.45万
-
财政年份:2009
-
负责人:Jianbo Shi
-
依托单位:
CAREER: Learning to See - A Unified Segmentation and Recognition Approach
-
批准号:0447953
-
项目类别:Standard Grant
-
资助金额:$40.0万
-
财政年份:2005
-
负责人:Jianbo Shi
-
依托单位:
RR:MACNet: Mobile Ad-hoc Camera Networks
-
批准号:0423891
-
项目类别:Continuing Grant
-
资助金额:$19.86万
-
财政年份:2004
-
负责人:Jianbo Shi
-
依托单位:
海外基金