Recommending MeSH terms for annotating biomedical articles.

Recommending MeSH terms for annotating biomedical articles.
复制标题

DOI:
10.1136/amiajnl-2010-000055
复制
发表时间:
2011-09
期刊:
Journal of the American Medical Informatics Association : JAMIA
影响因子:
--
通讯作者:
Lu Z
Lu Z
中科院分区:
其他
文献类型:
--
作者:
Huang M;Névéol A;Lu Z

文献摘要

参考文献

被引文献

相似文献

由于从科学文献中人工策展关键方面的成本很高,因此非常需要用于辅助该过程的自动化方法。在这里,我们报告了一种新的方法,以促进MeSH索引,一个具有挑战性的任务,分配MeSH术语的MEDLINE引文的存档和检索。与以前的方法自动MeSH长期分配,我们重新制定的索引任务作为一个排名问题,使相关的MeSH标题排名高于那些不相关的。具体来说,对于每个文件,我们检索20个邻居文件,从邻居获得一个列表的MeSH主标题,并排名的MeSH主标题使用ListNet的学习排名算法。我们在200个文档上训练了我们的算法,并在之前使用的200个文档的基准集和1000个文档的更大数据集上进行了测试。在基准数据集上进行测试,我们的方法达到了0.390的精度,召回率为0.712,平均精度(MAP)为0.626。与现有技术相比,我们观察到MAP的统计学显著改善高达39%(p值<0.001)。在较大的文件集上也取得了类似的重大改进。实验结果表明,我们的方法可以做出迄今为止最准确的MeSH预测,这表明它在对MeSH索引产生实际影响方面具有巨大潜力。此外,如所讨论的,所提出的学习框架是鲁棒的,并且可以适用于生物医学领域中MeSH索引之外的许多其他类似任务。所有数据集可在http://www.ncbi.nlm.nih.gov/CBBresearch/Lu/indexing上查阅。
Due to the high cost of manual curation of key aspects from the scientific literature, automated methods for assisting this process are greatly desired. Here, we report a novel approach to facilitate MeSH indexing, a challenging task of assigning MeSH terms to MEDLINE citations for their archiving and retrieval. Unlike previous methods for automatic MeSH term assignment, we reformulate the indexing task as a ranking problem such that relevant MeSH headings are ranked higher than those irrelevant ones. Specifically, for each document we retrieve 20 neighbor documents, obtain a list of MeSH main headings from neighbors, and rank the MeSH main headings using ListNet–a learning-to-rank algorithm. We trained our algorithm on 200 documents and tested on a previously used benchmark set of 200 documents and a larger dataset of 1000 documents. Tested on the benchmark dataset, our method achieved a precision of 0.390, recall of 0.712, and mean average precision (MAP) of 0.626. In comparison to the state of the art, we observe statistically significant improvements as large as 39% in MAP (p-value <0.001). Similar significant improvements were also obtained on the larger document set. Experimental results show that our approach makes the most accurate MeSH predictions to date, which suggests its great potential in making a practical impact on MeSH indexing. Furthermore, as discussed the proposed learning framework is robust and can be adapted to many other similar tasks beyond MeSH indexing in the biomedical domain. All data sets are available at: http://www.ncbi.nlm.nih.gov/CBBresearch/Lu/indexing.
DOI: 10.1093/jee/39.2.269
发表时间: 1945-01-01
期刊: BIOMETRICS BULLETIN
影响因子: --
作者:
WILCOXON, F
通讯作者: WILCOXON, F
DOI: 10.1093/bioinformatics/bti783
发表时间: 2006-03-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Ruch, P
通讯作者: Ruch, P
DOI: 10.1197/jamia.m2431
发表时间: 2008-07-01
影响因子: 6.4
作者:
Sohn, Sunghwan;Kim, Won;Wilbur, W. John
通讯作者: Wilbur, W. John
DOI: 10.1186/1471-2105-8-423
发表时间: 2007-10-30
期刊: BMC bioinformatics
影响因子: 3
作者:
Lin J;Wilbur WJ
通讯作者: Wilbur WJ
DOI: 10.1093/bioinformatics/bti503
发表时间: 2005-08-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Djebbari, A;Karamycheva, S;Quackenbush, J
通讯作者: Quackenbush, J