Extracting Semantics from Unconstrained Navigation on Wikipedia

Extracting Semantics from Unconstrained Navigation on Wikipedia
复制标题

DOI:
10.1007/s13218-015-0417-5
复制
发表时间:
2016-06
期刊:
KI - Künstliche Intelligenz
影响因子:
--
通讯作者:
Thomas Niebler;Daniel Schlör;Martin Becker;A. Hotho
Thomas Niebler;Daniel Schlör;Martin Becker;A. Hotho
中科院分区:
其他
文献类型:
--
作者:
Thomas Niebler;Daniel Schlör;Martin Becker;A. Hotho

文献摘要

被引文献

相似文献

单词之间的语义关联已经成功地从维基百科页面的导航中提取出来。然而,在相应的工作中使用的导航数据是稀疏的,并且预计会有偏差,因为它们是在游戏环境中收集的。在本文中,我们提出了这一限制,并探索是否也可以从无约束导航中提取语义相关性。为此,我们首先强调无约束导航和游戏数据之间的结构差异。然后,我们采用最先进的方法来提取维基百科路径上的语义相关性。我们将这种方法应用于来自两个无约束导航数据集的转换以及来自WikiGame的转换,并基于两个常见的黄金标准比较结果。通过比较无约束导航与WikiGame收集的路径,我们确认了预期的结构差异。与此结果一致的是,前面提到的用于导航数据语义提取的最先进的方法对于无约束导航不能产生良好的结果。然而,我们能够推导出一种既适用于无约束导航数据又适用于游戏数据的相关性度量。总的来说,我们表明维基百科上的无约束导航数据适合于提取语义。
Semantic relatedness between words has been successfully extracted from navigation on Wikipedia pages. However, the navigational data used in the corresponding works are sparse and expected to be biased since they have been collected in the context of games. In this paper, we raise this limitation and explore if semantic relatedness can also be extracted from unconstrained navigation. To this end, we first highlight structural differences between unconstrained navigation and game data. Then, we adapt a state of the art approach to extract semantic relatedness on Wikipedia paths. We apply this approach to transitions derived from two unconstrained navigation datasets as well as transitions from WikiGame and compare the results based on two common gold standards. We confirm expected structural differences when comparing unconstrained navigation with the paths collected by WikiGame. In line with this result, the mentioned state of the art approach for semantic extraction on navigation data does not yield good results for unconstrained navigation. Yet, we are able to derive a relatedness measure that performs well on both unconstrained navigation data as well as game data. Overall, we show that unconstrained navigation data on Wikipedia is suited for extracting semantics.