Interactive and context-aware tag spell check and correction

Interactive and context-aware tag spell check and correction
复制标题

交互式和上下文感知标签拼写检查和更正

DOI:
--
复制
发表时间:
2012
期刊:
International Conference on Information and Knowledge Management
影响因子:
--
通讯作者:
H. Vahabi
H. Vahabi
中科院分区:
--
文献类型:
--
作者:
F. Bonchi;O. Frieder;F. M. Nardini;F. Silvestri;H. Vahabi

文献摘要

被引文献

相似文献

协作内容创建和注释创建了各种媒体的巨大存储库,用户定义的标签扮演着核心角色,因为它们是组织,搜索和探索可用资源的简单而强大的工具。我们观察到,当用户使用一组标记来注释资源时,这些标记一次引入一个。因此,当引入第四标签时,由前三个标签表示的知识,即,产生第四标签的上下文是可用的,并且可用于产生当前标签的潜在校正。这个上下文,连同由标签在存储库的所有资源中的共同出现所表示的“群体的智慧”,可以被利用来提供交互式标签拼写检查和校正。我们开发这个想法的框架,基于加权标签同现图和加权邻域上定义的节点相关性措施。我们在来自YouTube的数据集上测试我们的建议。结果表明,我们的框架是有效的,因为它优于两个重要的基线。我们还表明,它是有效的,从而使其在现代标记服务中的使用。
Collaborative content creation and annotation creates vast repositories of all sorts of media, and user-defined tags play a central role as they are a simple yet powerful tool for organizing, searching and exploring the available resources. We observe that when a user annotates a resource with a set of tags, those tags are introduced one at a time. Therefore, when the fourth tag is introduced, a knowledge represented by the previous three tags, i.e., the context in which the fourth tag is produced, is available and exploitable for generating potential correction of the current tag. This context, together with the "wisdom of the crowd" represented by the co-occurrences of tags in all the resources of the repository, can be exploited to provide interactive tag spell check and correction. We develop this idea in a framework, based on a weighted tag co-occurrence graph and on nodes relatedness measures defined on weighted neighborhoods. We test our proposal on a dataset coming from YouTube. The results show that our framework is effective as it outperforms two important baselines. We also show that it is efficient, thus enabling its use in modern tagging services.