Segmentation-free compositional n-gram embedding

Segmentation-free compositional n-gram embedding
复制标题

DOI:
10.18653/v1/n19-1324
复制
发表时间:
2018-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Geewook Kim;Kazuki Fukui;Hidetoshi Shimodaira
Geewook Kim;Kazuki Fukui;Hidetoshi Shimodaira
中科院分区:
其他
文献类型:
--
作者:
Geewook Kim;Kazuki Fukui;Hidetoshi Shimodaira

文献摘要

相似文献

我们提出了一种新型的表示学习方法,无缝地建模单词,短语和句子。我们的方法不依赖于分词和任何人工注释的资源(例如,词典),但它是非常有效的噪音语料库写的未分割的语言,如中文和日语。我们方法的主要思想是完全忽略单词边界(即,无分段),并利用合成子n元语法的嵌入来构造原始语料库中的所有字符n元语法的表示。虽然这个想法很简单,但我们在各种基准测试和真实数据集上的实验表明了我们的建议的有效性。
We propose a new type of representation learning method that models words, phrases and sentences seamlessly. Our method does not depend on word segmentation and any human-annotated resources (e.g., word dictionaries), yet it is very effective for noisy corpora written in unsegmented languages such as Chinese and Japanese. The main idea of our method is to ignore word boundaries completely (i.e., segmentation-free), and construct representations for all character n-grams in a raw corpus with embeddings of compositional sub-n-grams. Although the idea is simple, our experiments on various benchmarks and real-world datasets show the efficacy of our proposal.