BioReddit: Word Embeddings for User-Generated Biomedical NLP

BioReddit: Word Embeddings for User-Generated Biomedical NLP
复制标题

DOI:
10.18653/v1/d19-6205
复制
发表时间:
2019-11
期刊:
--
影响因子:
--
通讯作者:
Marco Basaldella;Nigel Collier
Marco Basaldella;Nigel Collier
中科院分区:
其他
文献类型:
--
作者:
Marco Basaldella;Nigel Collier

文献摘要

相似文献

在过去的几年里,词嵌入以其不同的形状和迭代改变了自然语言处理的研究格局。生物医学文本处理领域对这场革命并不陌生;然而,该领域的学者主要是在科学文档上进行嵌入培训,即使是在处理用户生成的数据时也是如此。在本文中,我们展示了从医学论坛用户生成的文本收集的语料库中的训练嵌入如何严重影响下游任务的性能,当应用于用户生成的内容时,其性能优于在通用数据或科学论文上训练的嵌入。
Word embeddings, in their different shapes and iterations, have changed the natural language processing research landscape in the last years. The biomedical text processing field is no stranger to this revolution; however, scholars in the field largely trained their embeddings on scientific documents only, even when working on user-generated data. In this paper we show how training embeddings from a corpus collected from user-generated text from medical forums heavily influences the performance on downstream tasks, outperforming embeddings trained both on general purpose data or on scientific papers when applied on user-generated content.