Person-centric Story Understanding
Person-centric Story Understanding
批准号:
2763736
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2020
资助国家:
英国
项目状态:
已结题
起止时间:
2020 至 --
中文摘要
摘要:本项目旨在进一步从视频中对叙事进行个性化的、以人为本的理解。该项目属于EPSRC信息和通信技术研究领域。故事理解是一个复杂的问题,它不仅依赖于短期的低层次知识(如当前执行的操作),还依赖于预测更高层次语义信息(如人类身份、动机和意图)的能力。故事理解领域的大多数作品都是基于时间信息聚合信息——也就是说,它们会暂时总结视频片段,并试图通过对这些时间总结进行推理来解决与角色相关的任务。相比之下,我们的目标是根据角色而不是时间来总结视频。能够进行这种推理的系统的用途是多方面的:从通过自动生成每天上传的许多视频中出现的每个人的叙述来提高各种视频平台的可访问性,到通过检索特定于角色的视频片段来提高生产力。作为我们工作的一部分,我们开发了一种基于神经网络的计算模型,能够检索和浓缩特定个体的信息。这与关注对此类信息的一般(非个人)理解的流行方法是正交的。例如,大多数现代计算机视觉方法(如智能手机图片库中的方法)能够在手机上找到人们滑雪的图像或视频,或者能够识别特定人物的所有图像,但不能同时识别两者。然而,我们的方法不仅可以同时找到人们滑雪的视频,还可以找到手机主人滑雪的特定视频。我们不仅通过训练模型识别数据中出现的每个人来实现这一点,而且还允许它从外部信息源推断身份,例如从人的图像中。换句话说,即使我们的模型不能识别手机主人的名字,一张图片就足以找到相关的图片或视频。为了做到这一点,我们设计了自己的评估数据集,可用于评估该领域未来开发的任何方法。该研究已在2022年英国机器视觉会议上发表,并获得口头报告。为了进一步挑战我们的方法,我们打算将它们扩展到更长的时间窗口,并最终扩展到电影的整个情节线,这将需要设计一套全新的算法,并提出一个巨大的技术挑战。我们还计划通过开发我们自己的任务和数据集来进一步以人为中心的视频故事理解,因为目前还没有基准数据集。
英文摘要
Summary: The aim of this project is to further personalised human-centric understanding of narrative from videos. This project falls within the EPSRC information and communication technologies research area. Story understanding is a complex problem, relying not only on the short-term low-level knowledge (e.g. action performed right now), but also on the ability to predict higher-level semantic information such as human identities, motivations, and intents. Most works in the realm of story-understanding aggregate information based on temporal information -- that is, they summarise video segments temporally, and attempt to solve character-related tasks by reasoning over these temporal summaries. In contrast, our objective is to summarise the videos based on character, rather than time. The uses for a system capable of such reasoning are many-fold: from improving accessibility of various video platforms by automatically generating the narrative for each person appearing in one of the many videos uploaded each day, to increasing productivity by enabling retrieval of video segments specific to characters. As a part of our work, we have developed a neural-network-based computational model that is able to retrieve and condense information for a specific individual. This is orthogonal to prevailing methods that focus on generic (non-personal) understanding of such information. For example, most modern computer-vision methods such as ones in smartphone photo libraries are able to find images or videos on your phone where people are skiing, or are able to recognise all images of a particular person, but not both at the same time. Our method, however, can simultaneously find not just the videos where people are skiing, but also the particular videos where the phone owner is skiing. We achieve this by not only training the model to recognise every person appearing in the data, but also allowing it to infer identities from outside sources of information, for example from images of the person. In other words, even if our model does not recognise the phone owner by their name, a single image is enough to find relevant images or videos. In order to do so, we designed our own evaluation dataset that can be used to evaluate any future methods developed in this area. This research has been published at British Machine Vision Conference 2022 and has been awarded an oral presentation.In order to challenge our methods further, we intend to extend them to longer temporal windows and eventually entire plot lines of movies which will require designing a whole new set of algorithms and presents an enormous technological challenge. We also plan to further the human-centric story understanding in videos by developing our own tasks and datasets, as no benchmark datasets exist at present.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
基于CCN的新互联网架构体系对比分析及其路由缓冲策略研究
-
批准号:61103027
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2011
-
负责人:雷凯
-
依托单位:
网格中以情境为中心的应用自动化研究
-
批准号:60703054
-
项目类别:青年科学基金项目
-
资助金额:21.0万元
-
批准年份:2007
-
负责人:黄震春
-
依托单位: