A Study of Methods for the Generation of Domain-Aware Word Embeddings

A Study of Methods for the Generation of Domain-Aware Word Embeddings
复制标题

DOI:
10.1145/3397271.3401287
复制
发表时间:
2020-07
期刊:
Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval
影响因子:
--
通讯作者:
Dominic Seyler;Chengxiang Zhai
Dominic Seyler;Chengxiang Zhai
中科院分区:
其他
文献类型:
--
作者:
Dominic Seyler;Chengxiang Zhai

文献摘要

相似文献

词嵌入是许多文本数据应用程序的重要组件。在大多数工作中,使用在一般文本语料库上训练的“开箱即用”嵌入,但当应用于特定领域的设置时,它们可能不太有效。因此,如何创建“领域感知”的词嵌入是一个有趣的开放研究问题。在本文中,我们研究了三种方法来创建领域感知的词嵌入的基础上,一般和特定领域的文本语料库,包括拼接的嵌入向量,加权融合的文本数据,对齐的嵌入向量的插值。尽管所研究的策略是为特定领域的任务量身定制的,但它们足够普遍,可以应用于任何领域,而不是特定于单个任务。实验结果表明,三种方法都能很好地工作,但插值方法的效果始终最好。
Word embeddings are essential components for many text data applications. In most work, "out-of-the-box" embeddings trained on general text corpora are used, but they can be less effective when applied to domain-specific settings. Thus, how to create "domain-aware" word embeddings is an interesting open research question. In this paper, we study three methods for creating domain-aware word embeddings based on both general and domain-specific text corpora, including concatenation of embedding vectors, weighted fusion of text data, and interpolation of aligned embedding vectors. Even though the investigated strategies are tailored for domain-specific tasks, they are general enough to be applied to any domain and are not specific to a single task. Experimental results show that all three methods can work well, however, the interpolation method consistently works best.