Word and Sentence Segmentation in German: Overcoming Idiosyncrasies in the Use of Punctuation in Private Communication

Word and Sentence Segmentation in German: Overcoming Idiosyncrasies in the Use of Punctuation in Private Communication
复制标题

DOI:
10.1007/978-3-319-73706-5_6
复制
发表时间:
2017-09
期刊:
--
影响因子:
--
通讯作者:
Kyoko Sugisaki
Kyoko Sugisaki
中科院分区:
其他
文献类型:
--
作者:
Kyoko Sugisaki

文献摘要

相似文献

在本文中,我们提出了一种德语文本分割系统。我们将条件随机场(CRF)(一种统计序列模型)应用于私人通信中使用的一种文本类型。我们表明,通过分割单个标点符号,考虑独立行以及使用无监督的单词表示(即 Brown 聚类、Word2Vec 和 Fasttext),在私人通信中使用的明信片语料库中实现了 96% 的标签准确率。
In this paper, we present a segmentation system for German texts. We apply conditional random fields (CRF), a statistical sequential model, to a type of text used in private communication. We show that by segmenting individual punctuation, and by taking into account freestanding lines and that using unsupervised word representation (i. e., Brown clustering, Word2Vec and Fasttext) achieved a label accuracy of 96% in a corpus of postcards used in private communication.