MeSHLabeler: improving the accuracy of large-scale MeSH indexing by integrating diverse evidence.

MeSHLabeler: improving the accuracy of large-scale MeSH indexing by integrating diverse evidence.
复制标题

MeSHLabeler:通过整合不同证据提高大规模 MeSH 索引的准确性

DOI:
10.1093/bioinformatics/btv237
复制
发表时间:
2015-06-15
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Zhu S
Zhu S
中科院分区:
其他
文献类型:
--
作者:
Liu K;Peng S;Wu J;Zhai C;Mamitsuka H;Zhu S

文献摘要

被引文献

相似文献

动机:美国国家医学图书馆(NLM)使用医学主题词(MeSH)来索引MEDLINE中几乎所有的引文,这极大地方便了生物医学信息检索和文本挖掘的应用。为了减少手动注释的时间和财务成本,NLM 开发了一个软件包——医学文本索引器(MTI),用于辅助 MeSH 注释,该软件包使用 k 最近邻(KNN)、模式匹配和索引规则。其他类型的信息,例如 MeSH 分类器(单独训练)的预测,也可以用于自动 MeSH 注释。然而,现有方法无法有效整合多个证据进行 MeSH 注释。方法:我们提出了一个新颖的框架 MeSHLabeler,通过使用“学习排序”来整合多个证据以实现准确的 MeSH 注释。证据包括来自 MeSH 分类器、KNN、模式匹配、MTI 以及不同 MeSH 术语之间的相关性等的大量预测。每个 MeSH 分类器都是独立训练的,因此不同分类器的预测分数是不可比较的。为了解决这个问题,我们开发了一种有效的分数标准化程序来提高预测准确性。结果:MeSHLabeler 在 2014 年 BioASQ 挑战赛的任务 2A 中获得第一名,BioASQ 挑战赛提供的 9,040 次引用的 Micro F 测量值为 0.6248。请注意,该准确度比 MTI 获得的 0.5724 高出约 9.15%。可用性和实施​​:该软件可根据要求提供。联系方式:zhusf@fudan.edu.cn
Motivation: Medical Subject Headings (MeSHs) are used by National Library of Medicine (NLM) to index almost all citations in MEDLINE, which greatly facilitates the applications of biomedical information retrieval and text mining. To reduce the time and financial cost of manual annotation, NLM has developed a software package, Medical Text Indexer (MTI), for assisting MeSH annotation, which uses k-nearest neighbors (KNN), pattern matching and indexing rules. Other types of information, such as prediction by MeSH classifiers (trained separately), can also be used for automatic MeSH annotation. However, existing methods cannot effectively integrate multiple evidence for MeSH annotation. Methods: We propose a novel framework, MeSHLabeler, to integrate multiple evidence for accurate MeSH annotation by using ‘learning to rank’. Evidence includes numerous predictions from MeSH classifiers, KNN, pattern matching, MTI and the correlation between different MeSH terms, etc. Each MeSH classifier is trained independently, and thus prediction scores from different classifiers are incomparable. To address this issue, we have developed an effective score normalization procedure to improve the prediction accuracy. Results: MeSHLabeler won the first place in Task 2A of 2014 BioASQ challenge, achieving the Micro F-measure of 0.6248 for 9,040 citations provided by the BioASQ challenge. Note that this accuracy is around 9.15% higher than 0.5724, obtained by MTI. Availability and implementation: The software is available upon request. Contact: zhusf@fudan.edu.cn