Detailed and Deep Image Understanding
Detailed and Deep Image Understanding
批准号:
EP/L024683/1
负责人:
Andrea Vedaldi
金额:
$12.66万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2014
资助国家:
英国
项目状态:
已结题
起止时间:
2014 至 --
中文摘要
计算机视觉是一种允许机器自动理解图像内容的技术,正在推动数字图像处理的革命。例如,现在可以使用计算机在互联网上搜索数十亿张图像和数百万小时的视频,以获得特定内容(谷歌谷歌),解释手势和身体动作玩游戏(微软Kinect),自动将摄像头聚焦在人脸上,或者构建可以24小时监控危险工业设备的智能摄像头。如果不是因为它们的规模,这些任务对人类来说是微不足道的。然而,视觉在计算上是非常具有挑战性的,以至于我们大脑的一半以上都致力于这项功能。由于这种复杂性无法通过手工制作软件来满足,因此视觉架构现在可以利用先进的机器学习和优化技术,从数百万个示例图像中自动学习。然而,尽管最近取得了巨大的成功,但与人类的视觉相比,机器视觉仍然相形见绌。也许最令人失望的限制是,这些系统一次只能处理一个任务,比如判断一个特定的图像是否包含人。识别不同的概念,例如狗,或解决不同的任务,例如概述而不是识别人,需要从头开始学习一个新的系统,浪费时间和精力。我的研究想法是将现有架构转换为“视觉知识”的存储库,可以重用和增量扩展以解决多个任务和领域,大大提高效率,可扩展性,技术的灵活性。关键的科学挑战是了解视觉信息如何在最先进的视觉系统中编码。事实上,由于这些是自动学习的,而不是手工制作的,目前还不清楚它们捕获了什么信息以及如何表示。深入的调查将正式和定量地阐明这一点,并将成为越来越多的概念和任务之间共享和整合视觉知识的基础,包括系统初始设计中未涉及的概念和任务。与此同时,识别细粒度信息将使系统能够对图像内容进行更详细、更全面和更有意义的理解。由于拟议的研究将增强核心计算机视觉技术,该技术已经为无数应用提供了动力,因此其影响潜力巨大。例如,计算机将能够通过匹配使用更丰富的视觉词汇表达的更详细的查询来搜索图像;软件将能够以最小的努力扩展到新的领域和任务;计算机视觉系统将能够以明确直观的方式解释它们如何理解图像。研究成果将以最严格的方式在国际基准数据和协议上进行评估。研究结果将通过分发实施新技术的开放源码软件,向广大技术受众提供。该项目还可能产生强大的学术影响,巩固英国在计算机视觉(数字经济的战略竞争领域)的领导地位。
英文摘要
Computer vision, the technology that allows machines to understand the content of image automatically, is fuelling a revolution in digital image processing. For example, it is now possible to use computers to search billions of images and millions of hours of video in the Internet for a particular content (Google Googles), interpret gestures and body motions to play games (Microsoft Kinect), automatically focus cameras on faces, or build smart cameras that can monitor hazardous industrial equipment on a 24h basis.If not for their scale, these tasks would appear trivial to a human. However, vision is computationally exceptionally challenging, to the extent that more than half of our brain is dedicated to this function alone. Since this complexity cannot be met by hand-crafting software, vision architectures are nowadays learned automatically from million of example images, leveraging advanced machine learning and optimisation technologies. Despite recent terrific successes, however, machine vision still pales in comparison to vision in humans. Probably the most disappointing restriction is that these systems can address a single task at a time, such as deciding whether a particular image contains, say, person. Recognising a different concept, for example a dog, or addressing a different task, for example outlining rather than recognising the person, requires learning a new system from scratch, wasting time and effort.My research idea is to transform existing architectures into repositories of 'visual knowledge' that can be reused and extended incrementally to address multiple tasks and domains, greatly improving the efficiency, scalability, and flexibility of the technology. The key scientific challenge is to understand how visual information is encoded in state-of-the-art vision systems. In fact, since these are learned automatically rather than being hand-crafted, it is currently unclear what information is captured by them and how it is represented. An in-depth investigation will explicate this formally and quantitatively and will be the basis to share and integrate visual knowledge between a growing number of concepts and tasks, including ones not addressed by the initial design of the system. At the same time, identifying fine-grained information will allow a system to obtain a more detailed, comprehensive, and meaningful understanding of the content of images.The potential for impact is huge as the proposed research will enhance core computer vision technology that already powers countless applications. For example, computers will be able to search images by matching more detailed queries expressed using a far richer visual vocabulary; software will be extensible to new domains and tasks with minimal effort; and computer vision systems will be able to explain in explicit, intuitive terms how they understand images.The research outcomes will be evaluated in the most rigorous manner on international benchmark data and protocols. Research results will be made available to a widespread technical audience by distributing open source software implementing the new technology. The project is also likely to have a strong academic impact, consolidating the leadership of the UK in computer vision, a strategic competitive area in the digital economy.
期刊论文(5)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
登录
查看更多内容
Deep Seek引导下预防肝硬化腹水患者发生腹腔感染的约翰霍普金斯循证实践模型下中医护理策略的构建研究
-
批准号:2026JJ81909
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:胡曦
-
依托单位:
基于Deep Unrolling的高分辨近红外二区荧光分子断层成像方法研究
-
批准号:12271434
-
项目类别:面上项目
-
资助金额:46万元
-
批准年份:2022
-
负责人:贺小伟
-
依托单位:
基于深度森林(Deep Forest)模型的表面增强拉曼光谱分析方法研究
-
批准号:2020A151501709
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2020
-
负责人:谢怡
-
依托单位:
面向Deep Web的数据整合关键技术研究
-
批准号:61872168
-
项目类别:面上项目
-
资助金额:62.0万元
-
批准年份:2018
-
负责人:董永权
-
依托单位:
基于Deep-learning的三江源区冰川监测动态识别技术研究
-
批准号:51769027
-
项目类别:地区科学基金项目
-
资助金额:38.0万元
-
批准年份:2017
-
负责人:张大奇
-
依托单位:
具有时序处理能力的Spiking-Deep Learning(脉冲深度学习)方法研究
-
批准号:61573081
-
项目类别:面上项目
-
资助金额:64.0万元
-
批准年份:2015
-
负责人:屈鸿
-
依托单位:
基于语义计算的海量Deep Web知识探索机制研究
-
批准号:61272411
-
项目类别:面上项目
-
资助金额:80.0万元
-
批准年份:2012
-
负责人:赵峰
-
依托单位:
Deep Web数据集成查询结果抽取与整合关键技术研究
-
批准号:61100167
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2011
-
负责人:董永权
-
依托单位:
面向Deep Web的大规模知识库自动构建方法研究
-
批准号:61170020
-
项目类别:面上项目
-
资助金额:57.0万元
-
批准年份:2011
-
负责人:崔志明
-
依托单位:
Deep Web敏感聚合信息保护方法研究
-
批准号:61003054
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2010
-
负责人:赵朋朋
-
依托单位: