Improved biomedical word embeddings in the transformer era.
Improved biomedical word embeddings in the transformer era.
复制标题
DOI:
10.1016/j.jbi.2021.103867
复制
发表时间:
2021-08
影响因子:
4.5
通讯作者:
Kavuluru R
中科院分区:
文献类型:
--
作者:
Noh J;Kavuluru R
Recent natural language processing (NLP) research is dominated by neural network methods that employ word embeddings as basic building blocks. Pre-training with neural methods that capture local and global distributional properties (e.g., skip-gram, GLoVE) using free text corpora is often used to embed both words and concepts. Pre-trained embeddings are typically leveraged in downstream tasks using various neural architectures that are designed to optimize task-specific objectives that might further tune such embeddings. Despite advances in contextualized language model based embeddings, static word embeddings still form an essential starting point in BioNLP research and applications. They are useful in low resource settings and in lexical semantics studies. Our main goal is to build improved biomedical word embeddings and make them publicly available for downstream applications. We jointly learn word and concept embeddings by first using the skip-gram method and further fine-tuning them with correlational information manifesting in co-occurring Medical Subject Heading (MeSH) concepts in biomedical citations. This fine-tuning is accomplished with the transformer-based BERT architecture in the two-sentence input mode with a classification objective that captures MeSH pair co-occurrence. We conduct evaluations of these tuned static embeddings using multiple datasets for word relatedness developed by previous efforts. Both in qualitative and quantitative evaluations we demonstrate that our methods produce improved biomedical embeddings in comparison with other static embedding efforts. Without selectively culling concepts and terms (as was pursued by previous efforts), we believe we offer the most exhaustive evaluation of biomedical embeddings to date with clear performance improvements across the board. We repurposed a transformer architecture (typically used to generate dynamic embeddings) to improve static biomedical word embeddings using concept correlations. We provide our code and embeddings for public use for downstream applications and research endeavors: https://github.com/bionlproc/BERT-CRel-Embeddings
登录
查看更多内容
影响因子:
7.5
作者:
Henry, Sam;McQuilkin, Alex;McInnes, Bridget T.
通讯作者:
McInnes, Bridget T.
影响因子:
4.5
作者:
Pedersen, Ted;Pakhomov, Serguei V. S.;Chute, Christopher G.
通讯作者:
Chute, Christopher G.
DOI:
10.18653/v1/2020.findings-emnlp.304
发表时间:
2020-11
期刊:
Proceedings of the Conference on Empirical Methods in Natural Language Processing. Conference on Empirical Methods in Natural Language Processing
影响因子:
--
作者:
Noh J;Kavuluru R
通讯作者:
Kavuluru R
影响因子:
7.5
作者:
Kavuluru R;Rios A;Lu Y
通讯作者:
Lu Y
DOI:
10.1080/00437956.1954.11659520
发表时间:
1954-08-01
影响因子:
0.6
作者:
Harris, Zellig S.
通讯作者:
Harris, Zellig S.