Effects of information and machine learning algorithms on word sense disambiguation with small datasets

Effects of information and machine learning algorithms on word sense disambiguation with small datasets
复制标题

DOI:
10.1016/j.ijmedinf.2005.03.013
复制
发表时间:
2005-08-01
影响因子:
4.9
通讯作者:
Rindflesch, TC
Rindflesch, TC
中科院分区:
医学2区
文献类型:
--
作者:
Leroy, G;Rindflesch, TC

文献摘要

被引文献

相似文献

当前的词义消歧方法使用(并且通常联合收割机)各种机器学习技术。大多数是指歧义及其周围的词的特点,并基于成千上万的例子。不幸的是,开发大型训练集是繁重的,为了应对这一挑战,我们研究了在小型数据集上使用符号知识。一个朴素贝叶斯分类器被训练为15个单词,每个单词100个例子。统一医学语言系统(UMLS)的语义类型分配给句子中的概念和这些语义类型之间的关系形成的知识库。一个词最常见的意义作为基线。越来越准确的符号知识的效果进行了评估,在九个实验条件。性能通过基于10倍交叉验证的准确度来衡量。最佳条件是只使用句子中单词的语义类型。准确性平均比基线高10%;然而,它从8%恶化到29%改善不等。为了研究这种大的差异,我们进行了几次后续评估,测试了其他算法(决策树和神经网络)和黄金标准(每个专家),但结果没有显著差异。然而,我们注意到一个趋势,即最好的消歧是为那些对人类评估者来说最不麻烦的单词找到的。我们的结论是,无论是算法还是个人行为都不会导致这些巨大的差异,但是UMLS元词库的结构(用于表示歧义词的含义)会导致金标准的不准确性,从而导致词义消歧技术的不同表现。(C)2005爱思唯尔爱尔兰有限公司保留所有权利。
Current approaches to word sense disambiguation use (and often combine) various machine learning techniques. Most refer to characteristics of the ambiguity and its surrounding words and are based on thousands of examples. Unfortunately, developing large training sets is burdensome, and in response to this challenge, we investigate the use of symbolic knowledge for small datasets. A naive Bayes classifier was trained for 15 words with 100 examples for each. Unified Medical Language System (UMLS) semantic types assigned to concepts found in the sentence and relationships between these semantic types form the knowledge base. The most frequent sense of a word served as the baseline. The effect of increasingly accurate symbolic knowledge was evaluated in nine experimental conditions. Performance was measured by accuracy based on 10-fold cross-validation. The best condition used only the semantic types of the words in the sentence. Accuracy was then on average 10% higher than the baseline; however, it varied from 8% deterioration to 29% improvement. To investigate this large variance, we performed several follow-up evaluations, testing additional algorithms (decision tree and neural network), and gold standards (per expert), but the results did not significantly differ. However, we noted a trend that the best disambiguation was found for words that were the least troublesome to the human evaluators. We conclude that neither algorithm nor individual human behavior cause these large differences, but that the structure of the UMLS Metathesaurus (used to represent senses of ambiguous words) contributes to inaccuracies in the gold standard, leading to varied performance of word sense disambiguation techniques. (C) 2005 Elsevier Ireland Ltd. All rights reserved.