Crowdsourced Semantic Matching of Multi-Label Annotations

Crowdsourced Semantic Matching of Multi-Label Annotations
复制标题

DOI:
--
复制
发表时间:
2015-07
期刊:
--
影响因子:
--
通讯作者:
Lei Duan;S. Oyama;M. Kurihara;Haruhiko Sato
Lei Duan;S. Oyama;M. Kurihara;Haruhiko Sato
中科院分区:
其他
文献类型:
--
作者:
Lei Duan;S. Oyama;M. Kurihara;Haruhiko Sato

文献摘要

被引文献

相似文献

大多数多标签域缺乏权威的分类法。因此,在同一领域中通常使用不同的分类法,这导致了复杂性。虽然这种情况经常发生,但很少使用原则性统计方法对其进行研究。假设(1)在相同领域中使用的不同分类法通常建立在相同的潜在语义空间上,其中分类法中的每个可能的标签集表示单个语义概念,并且(2)众包在以低成本识别语义概念和实例之间的关系方面是有益的,我们提出了一种新的概率级联方法,用于在众包设置中建立语义匹配函数,该函数将标签集映射到一个(源)taxonomy可以根据它们之间的语义距离来标记另一个(目标)分类中的集合。已建立的函数可以用于直接从实例在源分类中的关联标签集检测实例在目标分类中的关联标签集,而无需任何额外的工作。在真实数据(叙事语句的情感标注)上的实验结果表明,该方法能够从有限数量的众包标注中稳健地建立表现出令人满意性能的语义匹配函数。
Most multi-label domains lack an authoritative taxonomy. Therefore, different taxonomies are commonly used in the same domain, which results in complications. Although this situation occurs frequently, there has been little study of it using a principled statistical approach. Given that (1) different taxonomies used in the same domain are generally founded on the same latent semantic space, where each possible label set in a taxonomy denotes a single semantic concept, and that (2) crowdsourcing is beneficial in identifying relationships between semantic concepts and instances at low cost, we proposed a novel probabilistic cascaded method for establishing a semantic matching function in a crowdsourcing setting that maps label sets in one (source) taxonomy to label sets in another (target) taxonomy in terms of the semantic distances between them. The established function can be used to detect the associated label set in the target taxonomy for an instance directly from its associated label set in the source taxonomy without any extra effort. Experimental results on real-world data (emotion annotations for narrative sentences) demonstrated that the proposed method can robustly establish semantic matching functions exhibiting satisfactory performance from a limited number of crowdsourced annotations.