Learning Image-Text Associations

Learning Image-Text Associations
复制标题

DOI:
10.1109/tkde.2008.150
复制
发表时间:
2009-02
影响因子:
8.9
通讯作者:
Tao Jiang;A. Tan
Tao Jiang;A. Tan
中科院分区:
计算机科学2区
文献类型:
--
作者:
Tao Jiang;A. Tan

文献摘要

被引文献

相似文献

Web信息融合可以定义为在万维网上整理和跟踪与特定主题相关的信息的问题。鉴于大多数现有的Web信息融合的工作都集中在基于文本的多文档摘要,本文关注的主题的图像和文本的关联,跨媒体Web信息融合的基石。具体来说,我们提出了两种基于小训练数据集发现图像和文本之间潜在关联的学习方法。第一种方法基于模糊变换,通过一组预定义的特定领域的信息类别来度量视觉特征和文本特征之间的信息相似性。另一种方法使用神经网络来通过自动地和递增地将相关联的特征汇总到一组信息模板中来学习视觉特征和文本特征之间的直接映射。尽管他们不同的方法,我们的实验结果恐怖分子域文档集显示,这两种方法都能够学习图像和文本之间的关联,从一个小的训练数据集。
Web information fusion can be defined as the problem of collating and tracking information related to specific topics on the World Wide Web. Whereas most existing work on Web information fusion has focused on text-based multidocument summarization, this paper concerns the topic of image and text association, a cornerstone of cross-media Web information fusion. Specifically, we present two learning methods for discovering the underlying associations between images and texts based on small training data sets. The first method based on vague transformation measures the information similarity between the visual features and the textual features through a set of predefined domain-specific information categories. Another method uses a neural network to learn direct mapping between the visual and textual features by automatically and incrementally summarizing the associated features into a set of information templates. Despite their distinct approaches, our experimental results on a terrorist domain document set show that both methods are capable of learning associations between images and texts from a small training data set.