Optimal training sets for Bayesian prediction of MeSH® assignment

Optimal training sets for Bayesian prediction of MeSH® assignment
复制标题

DOI:
10.1197/jamia.m2431
复制
发表时间:
2008-07-01
影响因子:
6.4
通讯作者:
Wilbur, W. John
Wilbur, W. John
中科院分区:
管理学2区
文献类型:
--
作者:
Sohn, Sunghwan;Kim, Won;Wilbur, W. John

文献摘要

被引文献

相似文献

目的:本研究的目的是提高朴素贝叶斯预测的医学主题词(MeSH)分配到documents.Design的主动学习启发的方法发现,使用最佳的训练集:作者选择了20 MeSH条款,其出现频率覆盖范围。对于每个MeSH术语,他们找到了一个最佳训练集,即整个训练集的子集。最佳训练集由包括给定MeSH术语(C-1类)的所有文档和不包括给定MeSH术语(C-1类)的最接近C-1类的那些文档组成。这些小集合用于预测MEDLINE(R)数据库中的MeSH分配。测量:使用在整个训练集、最佳集和随机集上训练的朴素贝叶斯学习器,使用平均精度来比较MeSH分配。作者比较了朴素贝叶斯平均精度的95%置信下限与K-最近邻(KNN)classification.Results的平均精度的上限:对于所有20个MeSH分配,最佳训练集产生了近200%的改善使用整个训练集。在其中17个MeSH任务中,使用最佳训练集的朴素贝叶斯在统计学上优于KNN。在其中的15个中,最佳训练集的表现优于优化的特征选择。总体而言,朴素贝叶斯平均比所有20个MeSH分配的KNN好14%。使用这些最佳集与另一个分类器,C-修改的最小二乘(CMLS),产生了额外的6%的改善超过朴素贝叶斯。结论:使用一个较小的最佳训练集大大提高了学习与朴素贝叶斯。性能优于KNN的上级。小训练集可以与其他复杂的学习方法一起使用,例如CMLS,其中使用整个训练集是不可行的。
Objectives: The aim of this study was to improve naive Bayes prediction of Medical Subject Headings (MeSH) assignment to documents using optimal training sets found by an active learning inspired method.Design: The authors selected 20 MeSH terms whose occurrences cover a range of frequencies. For each MeSH term, they found an optimal training set, a subset of the whole training set. An optimal training set consists of all documents including a given MeSH term (C-1 class) and those documents not including a given MeSH term (C-1 class) that are closest to the C-1 class. These small sets were used to predict MeSH assignments in the MEDLINE (R) database.Measurements: Average precision was used to compare MeSH assignment using the naive Bayes learner trained on the whole training set, optimal sets, and random sets. The authors compared 95% lower confidence limits of average precisions of naive Bayes with upper bounds for average precisions of a K-nearest neighbor (KNN) classifier.Results: For all 20 MeSH assignments, the optimal training sets produced nearly 200% improvement over use of the whole training sets. In 17 of those MeSH assignments, naive Bayes using optimal training sets was statistically better than a KNN. In 15 of those, optimal training sets performed better than optimized feature selection. Overall naive Bayes averaged 14% better than a KNN for all 20 MeSH assignments. Using these optimal sets with another classifier, C-modified least squares (CMLS), produced an additional 6% improvement over naive Bayes.Conclusion: Using a smaller optimal training set greatly improved learning with naive Bayes. The performance is superior to a KNN. The small training set can be used with other sophisticated learning methods, such as CMLS, where using the whole training set would not be feasible.