Improved biomedical word embeddings in the transformer era.

Improved biomedical word embeddings in the transformer era.
复制标题

DOI:
10.1016/j.jbi.2021.103867
复制
发表时间:
2021-08
影响因子:
4.5
通讯作者:
Kavuluru R
Kavuluru R
中科院分区:
医学3区
文献类型:
--
作者:
Noh J;Kavuluru R

文献摘要

参考文献

被引文献

相似文献

最近的自然语言处理(NLP)研究被神经网络方法所主导,这些方法使用单词嵌入作为基本的构建块。使用自由文本语料库使用捕捉局部和全局分布属性(例如,跳过语法、手套)的神经方法的预训练通常被用于嵌入单词和概念。预先训练的嵌入通常在下游任务中使用各种神经体系结构来利用,这些神经体系结构被设计成优化可能进一步调整这种嵌入的特定任务目标。尽管在基于上下文语言模型的嵌入方面取得了进展,但静态单词嵌入仍然是BioNLP研究和应用的重要起点。它们在低资源环境和词汇语义研究中很有用。我们的主要目标是构建改进的生物医学单词嵌入,并使其公开可用于下游应用。我们共同学习单词和概念的嵌入,首先使用跳过语法方法,然后用生物医学引文中共出现的医学主题标题(MESH)概念中体现的相关信息进一步微调它们。这种微调是通过基于变压器的BERT架构在两句话输入模式下完成的,其分类目标是捕获网状对的共现。我们使用先前努力开发的多个数据集对这些调优的静态嵌入进行了评估。在定性和定量评估中,我们都证明了与其他静态嵌入工作相比,我们的方法产生了更好的生物医学嵌入。没有选择性地挑选概念和术语(就像之前的努力所追求的那样),我们相信我们提供了迄今为止最详尽的生物医学嵌入评估,并在所有方面都有明显的性能改进。我们重新调整了转换器体系结构的用途(通常用于生成动态嵌入),以使用概念关联来改进静态生物医学单词嵌入。我们提供我们的代码和嵌入,供公众用于下游应用程序和研究工作:https://github.com/bionlproc/BERT-CRel-Embeddings
Recent natural language processing (NLP) research is dominated by neural network methods that employ word embeddings as basic building blocks. Pre-training with neural methods that capture local and global distributional properties (e.g., skip-gram, GLoVE) using free text corpora is often used to embed both words and concepts. Pre-trained embeddings are typically leveraged in downstream tasks using various neural architectures that are designed to optimize task-specific objectives that might further tune such embeddings. Despite advances in contextualized language model based embeddings, static word embeddings still form an essential starting point in BioNLP research and applications. They are useful in low resource settings and in lexical semantics studies. Our main goal is to build improved biomedical word embeddings and make them publicly available for downstream applications. We jointly learn word and concept embeddings by first using the skip-gram method and further fine-tuning them with correlational information manifesting in co-occurring Medical Subject Heading (MeSH) concepts in biomedical citations. This fine-tuning is accomplished with the transformer-based BERT architecture in the two-sentence input mode with a classification objective that captures MeSH pair co-occurrence. We conduct evaluations of these tuned static embeddings using multiple datasets for word relatedness developed by previous efforts. Both in qualitative and quantitative evaluations we demonstrate that our methods produce improved biomedical embeddings in comparison with other static embedding efforts. Without selectively culling concepts and terms (as was pursued by previous efforts), we believe we offer the most exhaustive evaluation of biomedical embeddings to date with clear performance improvements across the board. We repurposed a transformer architecture (typically used to generate dynamic embeddings) to improve static biomedical word embeddings using concept correlations. We provide our code and embeddings for public use for downstream applications and research endeavors: https://github.com/bionlproc/BERT-CRel-Embeddings
DOI: 10.1016/j.artmed.2018.08.006
发表时间: 2019-01-01
影响因子: 7.5
作者:
Henry, Sam;McQuilkin, Alex;McInnes, Bridget T.
通讯作者: McInnes, Bridget T.
DOI: 10.1016/j.jbi.2006.06.004
发表时间: 2007-06-01
影响因子: 4.5
作者:
Pedersen, Ted;Pakhomov, Serguei V. S.;Chute, Christopher G.
通讯作者: Chute, Christopher G.
DOI: 10.18653/v1/2020.findings-emnlp.304
发表时间: 2020-11
期刊: Proceedings of the Conference on Empirical Methods in Natural Language Processing. Conference on Empirical Methods in Natural Language Processing
影响因子: --
作者:
Noh J;Kavuluru R
通讯作者: Kavuluru R
DOI: 10.1016/j.artmed.2015.04.007
发表时间: 2015-10
影响因子: 7.5
作者:
Kavuluru R;Rios A;Lu Y
通讯作者: Lu Y
DOI: 10.1080/00437956.1954.11659520
发表时间: 1954-08-01
影响因子: 0.6
作者:
Harris, Zellig S.
通讯作者: Harris, Zellig S.