FullMeSH: improving large-scale MeSH indexing with full text

FullMeSH: improving large-scale MeSH indexing with full text
复制标题

FullMeSH:利用全文改进大规模 MeSH 索引

DOI:
10.1093/bioinformatics/btz756
复制
发表时间:
2020-03-01
期刊:
影响因子:
5.8
通讯作者:
Zhu, Shanfeng
Zhu, Shanfeng
中科院分区:
生物学3区
文献类型:
--
作者:
Dai, Suyang;You, Ronghui;Zhu, Shanfeng

文献摘要

被引文献

相似文献

动机 随着生物医学文献的快速增长,利用医学主题词(Medical Subject Heading,MeSH)自动标引生物医学文献,即MeSH标引,对于促进假说生成和知识发现变得越来越重要。在过去的几年里,已经提出了许多大规模的MeSH索引方法,例如医学文本索引器(MTI),MeSHLabeler,DeepMeSH和MeSHProbeNet。然而,这些方法的性能受到限制,使用有限的信息,即只有标题和摘要的生物医学文章。 结果 我们提出了FullMeSH,一个大规模的MeSH索引方法,利用最近增加的全文文章的可用性。与DeepMeSH和其他最先进的方法相比,FullMeSH有三个新颖之处:1)FullMeSH不是将全文作为一个整体使用,而是将其分割成几个部分,并带有规范化的标题,以区分它们对整体性能的贡献。2)FullMeSH通过结合稀疏和深层语义表示,将来自不同部分的证据集成在“学习排名”框架中。3)FullMeSH为每个部分训练基于注意力的卷积神经网络(AttentionCNN),从而在不常见的MeSH标题上实现更好的性能。FullMeSH是在PubMed Central Open Access子集中的140万篇全文文章的整个集合上开发和经验培训的。它在10,000篇文章的测试集上实现了66.76%的Micro F测量,分别比DeepMeSH和MeSHLabeler高3.3%和6.4%。此外,FullMeSH在索引检查标签(一组最常索引的MeSH标题)方面比DeepMeSH平均提高了4.7%。 补充资料 补充数据可在Bioinformatics在线获得。
MOTIVATION With the rapidly growing biomedical literature, automatically indexing biomedical articles by Medical Subject Heading (MeSH), namely MeSH indexing, has become increasingly important for facilitating hypothesis generation and knowledge discovery. Over the past years, many large-scale MeSH indexing approaches have been proposed, such as Medical Text Indexer (MTI), MeSHLabeler, DeepMeSH and MeSHProbeNet. However, the performance of these methods is hampered by using limited information, i.e. only the title and abstract of biomedical articles. RESULTS We propose FullMeSH, a large-scale MeSH indexing method taking advantage of the recent increase in the availability of full text articles. Compared to DeepMeSH and other state-of-the-art methods, FullMeSH has three novelties: 1) Instead of using a full text as a whole, FullMeSH segments it into several sections with their normalized titles in order to distinguish their contributions to the overall performance. 2) FullMeSH integrates the evidence from different sections in a "learning to rank" framework by combining the sparse and deep semantic representations. 3) FullMeSH trains an Attention-based Convolutional Neural Network (AttentionCNN) for each section, which achieves better performance on infrequent MeSH headings. FullMeSH has been developed and empirically trained on the entire set of 1.4 million full-text articles in the PubMed Central Open Access subset. It achieved a Micro F-measure of 66.76% on a test set of 10,000 articles, which was 3.3% and 6.4% higher than DeepMeSH and MeSHLabeler, respectively. Furthermore, FullMeSH demonstrated an average improvement of 4.7% over DeepMeSH for indexing Check Tags, a set of most frequently indexed MeSH headings. SUPPLEMENTARY INFORMATION Supplementary data are available at Bioinformatics online.