Towards robust word embeddings for noisy texts

Towards robust word embeddings for noisy texts
复制标题

针对嘈杂文本的稳健词嵌入

DOI:
10.3390/app10196893
复制
发表时间:
2019
期刊:
ArXiv
影响因子:
--
通讯作者:
Carlos Gómez
Carlos Gómez
中科院分区:
--
文献类型:
--
作者:
Yerai Doval;Jesús Vilares;Carlos Gómez

文献摘要

被引文献

相似文献

对词嵌入的研究主要集中在提高它们在标准语料库上的性能,而忽略了社交媒体上以推文和其他类型的非标准写作形式出现的嘈杂文本所带来的困难。在这项工作中,我们提出了一个简单的扩展skipgram模型中,我们引入了桥字的概念,这是人工的话添加到模型中,以加强标准的话和他们的噪声变体之间的相似性。我们的新嵌入在各种内在和外在的评估任务上都优于噪声文本的基线模型,同时在标准文本上保持良好的性能。据我们所知,这是第一个在单词嵌入级别处理这些类型的噪声文本的显式方法,它超越了对词汇表外单词的支持。
Research on word embeddings has mainly focused on improving their performance on standard corpora, disregarding the difficulties posed by noisy texts in the form of tweets and other types of non-standard writing from social media. In this work, we propose a simple extension to the skipgram model in which we introduce the concept of bridge-words, which are artificial words added to the model to strengthen the similarity between standard words and their noisy variants. Our new embeddings outperform baseline models on noisy texts on a wide range of evaluation tasks, both intrinsic and extrinsic, while retaining a good performance on standard texts. To the best of our knowledge, this is the first explicit approach at dealing with these types of noisy texts at the word embedding level that goes beyond the support for out-of-vocabulary words.