BioWordVec, improving biomedical word embeddings with subword information and MeSH

BioWordVec, improving biomedical word embeddings with subword information and MeSH
复制标题

DOI:
10.1038/s41597-019-0055-0
复制
发表时间:
2019-05-10
期刊:
影响因子:
9.8
通讯作者:
Lu, Zhiyong
Lu, Zhiyong
中科院分区:
综合性期刊2区
文献类型:
--
作者:
Zhang, Yijia;Chen, Qingyu;Lu, Zhiyong

文献摘要

被引文献

相似文献

分布式词语表示已经成为生物医学自然语言处理、文本挖掘和信息检索的重要基础。传统上,词嵌入是从大量未标记文本在词级计算的,忽略了词的内部结构中存在的信息或领域特定结构化资源(例如本体)中可用的任何信息。然而,正如最近在一般领域的一些研究所表明的那样,这种信息具有极大地提高单词表征质量的潜力。在这里,我们介绍了BioWordVec:一组开放的生物医学词汇向量/嵌入,它将来自未标记的生物医学文本的子词信息与广泛使用的生物医学控制词汇表相结合,称为医学主题标题(MESH)。我们评估了我们在生物医学领域的多个NLP任务中生成的词嵌入的有效性和实用性。我们的基准测试结果表明,在这些具有挑战性的任务中,我们的单词嵌入可以导致比以前最先进的水平显著提高的性能。
Distributed word representations have become an essential foundation for biomedical natural language processing (BioNLP), text mining and information retrieval. Word embeddings are traditionally computed at the word level from a large corpus of unlabeled text, ignoring the information present in the internal structure of words or any information available in domain specific structured resources such as ontologies. However, such information holds potentials for greatly improving the quality of the word representation, as suggested in some recent studies in the general domain. Here we present BioWordVec: an open set of biomedical word vectors/embeddings that combines subword information from unlabeled biomedical text with a widely-used biomedical controlled vocabulary called Medical Subject Headings (MeSH). We assess both the validity and utility of our generated word embeddings over multiple NLP tasks in the biomedical domain. Our benchmarking results demonstrate that our word embeddings can result in significantly improved performance over the previous state of the art in those challenging tasks.