An active learning framework for content-based information retrieval

An active learning framework for content-based information retrieval
复制标题

DOI:
10.1109/tmm.2002.1017738
复制
发表时间:
2002-06-01
影响因子:
7.3
通讯作者:
Chen, TS
Chen, TS
中科院分区:
计算机科学1区
文献类型:
--
作者:
Zhang, C;Chen, TS

文献摘要

被引文献

相似文献

在本文中,我们提出了一个通用的主动学习框架,基于内容的信息检索(CBIR)。我们使用这个框架来指导隐藏的注释,以提高检索性能。对于数据库中的每个对象,我们维护一个概率列表,每个概率表示该对象具有其中一个属性的概率。在训练过程中,学习算法对数据库中的对象进行采样,并将它们呈现给注释器以分配属性。对于每个采样对象,每个概率都被设置为1或0,这取决于注释器是否分配了相应的属性。对于未注释的对象,学习算法使用有偏核回归来估计它们的概率。然后定义知识增益,以确定在未注释的对象中,系统最不确定的对象。然后,系统将其作为下一个样本呈现给注释器,并为其分配属性。在检索过程中,概率列表作为一个特征向量,用于计算两个对象之间的语义距离,或者用户查询与数据库中的对象之间的语义距离。两个对象之间的总体距离由语义距离和低级特征距离的加权和确定。该算法在合成数据库和真实的三维模型数据库上进行了测试。在这两种情况下,系统的检索性能随着注释样本的数量而迅速提高。此外,我们表明,主动学习优于基于随机抽样的学习。
In this paper, we propose a general active learning framework for content-based information retrieval (CBIR). We use this framework to guide hidden annotations in order to improve the retrieval performance. For each object in the database, we maintain a list of probabilities, each indicating the probability of this object having one of the attributes. During training, the learning algorithm samples objects in the database and presents them to the annotator to assign attributes to. For each sampled object, each probability is set to be one or zero depending on whether or not the corresponding attribute is assigned by the annotator. For objects that have not been annotated, the learning algorithm estimates their probabilities with biased kernel regression. Knowledge gain is then defined to determine, among the objects that have not been annotated, which one the system is the most uncertain of. The system then presents it as the next sample to the annotator to which it is assigned attributes. During retrieval, the list of probabilities works as a feature vector for us to calculate the semantic distance between two objects, or between the user query and an object in the database. The overall distance between two objects is determined by a weighted sum of the semantic distance and the low-level feature distance. The algorithm is tested on both synthetic databases and real databases of three-dimensional (3-D) models. In both cases, the retrieval performance of the system improves rapidly with the number of annotated samples. Furthermore, we show that active learning outperforms learning based on random sampling.