Word sense disambiguation for free-text indexing using a massive semantic network

Word sense disambiguation for free-text indexing using a massive semantic network
复制标题

DOI:
10.1145/170088.170106
复制
发表时间:
1993-12
期刊:
--
影响因子:
--
通讯作者:
Michael Sussna
Michael Sussna
中科院分区:
其他
文献类型:
--
作者:
Michael Sussna

文献摘要

被引文献

相似文献

无语义、基于词的信息检索受到两个互补问题的阻碍。首先,当搜索词的所有含义都被使用时,搜索相关文档会返回不相关的项,而不仅仅是预期的含义。这导致精度低。第二,当相关项目不是在实际搜索项下索引,而是在相关项下索引时,就会错过相关项目。这导致低召回。使用无语义的方法,通常无法同时提高查准率和查全率。在文献标引过程中进行词义消歧可以提高精度。我们研究了在索引过程中使用大规模的Word Net语义网络进行消歧。在SMART检索环境的无约束文本中,我们必须从输入文本中获得我们自己的内容描述,仅对输入进行词性标记。我们采用网络节点之间的语义距离的概念。通过从一组相邻术语中找到多个含义的组合来消除具有多个含义的输入文本术语的歧义,该组合最小化含义之间的总成对距离。到目前为止,结果令人鼓舞。与偶然性相比,
Semantics-free, word-based information retrieval is thwarted by two complementary problems. First, search for relevant documents returns irrelevant items when all meanings of a search term are used, rather than just the meaning intended. This causes low precision. Second, relevant items are missed when they are indexed not under the actual search terms, but rather under related terms. This causes low recall. With semantics-free approaches there is generally no way to improve both precision and recall at the same time. Word sense disambiguation during document indexing should improve precision. We have investigated using the massive Word Net semantic network for disambigu at ion during indexing. With the unconstrained text of the SMART ret rieval environment, we have had to derive our own content description from the input text, given only part-ofspeech tagging of the input. We employ the notion of semantic distance between network nodes. Input text terms with multiple senses are disambiguated by finding the combination of senses from a set of contiguous terms which minimizes total pairwise dist ante between senses. Results so far have been encouraging. Improvement in disamblguation compared with chance is clear