Learning Everything about Anything: Webly-Supervised Visual Concept Learning

Learning Everything about Anything: Webly-Supervised Visual Concept Learning
复制标题

DOI:
10.1109/cvpr.2014.412
复制
发表时间:
2014-06
期刊:
2014 IEEE Conference on Computer Vision and Pattern Recognition
影响因子:
--
通讯作者:
S. Divvala;Ali Farhadi;Carlos Guestrin
S. Divvala;Ali Farhadi;Carlos Guestrin
中科院分区:
其他
文献类型:
--
作者:
S. Divvala;Ali Farhadi;Carlos Guestrin

文献摘要

被引文献

相似文献

认可是从实验室到现实世界的应用。虽然看到它的潜力被挖掘是令人鼓舞的,但它给视觉研究人员带来了一个根本性的挑战:可扩展性。我们如何才能学习一个模型,用于任何概念,详尽地涵盖其所有外观变化,同时需要最少或根本不需要人类监督来编译视觉变化的词汇表,收集训练图像和注释,并学习模型?在本文中,我们介绍了一种全自动的方法,用于学习任何概念中各种变化(例如动作,交互,属性等)的广泛模型。我们的方法利用大量的在线书籍资源来发现方差的词汇表,并将数据收集和建模步骤交织在一起,以减轻在训练模型时对明确的人工监督的需求。我们的方法以一种方便和有用的方式组织关于概念的视觉知识,从而实现视觉和NLP的各种应用。我们的在线系统已经被用户查询,以学习几个有趣的概念,包括早餐,甘地,美丽的模型等,到目前为止,我们的系统有超过50000个变化的150个概念,并已注释超过1000万个图像与边界框。
Recognition is graduating from labs to real-world applications. While it is encouraging to see its potential being tapped, it brings forth a fundamental challenge to the vision researcher: scalability. How can we learn a model for any concept that exhaustively covers all its appearance variations, while requiring minimal or no human supervision for compiling the vocabulary of visual variance, gathering the training images and annotations, and learning the models? In this paper, we introduce a fully-automated approach for learning extensive models for a wide range of variations (e.g. actions, interactions, attributes and beyond) within any concept. Our approach leverages vast resources of online books to discover the vocabulary of variance, and intertwines the data collection and modeling steps to alleviate the need for explicit human supervision in training the models. Our approach organizes the visual knowledge about a concept in a convenient and useful way, enabling a variety of applications across vision and NLP. Our online system has been queried by users to learn models for several interesting concepts including breakfast, Gandhi, beautiful, etc. To date, our system has models available for over 50, 000 variations within 150 concepts, and has annotated more than 10 million images with bounding boxes.