Inducing Word and Part-of-Speech with Pitman-Yor Hidden Semi-Markov Models

Inducing Word and Part-of-Speech with Pitman-Yor Hidden Semi-Markov Models
复制标题

DOI:
10.3115/v1/p15-1171
复制
发表时间:
2015-07
期刊:
--
影响因子:
--
通讯作者:
Kei Uchiumi;Hiroshi Tsukahara;D. Mochihashi
Kei Uchiumi;Hiroshi Tsukahara;D. Mochihashi
中科院分区:
其他
文献类型:
--
作者:
Kei Uchiumi;Hiroshi Tsukahara;D. Mochihashi

文献摘要

相似文献

我们提出了一种非参数贝叶斯模型,用于联合无监督分词和原始字符串词性标注。我们的模型扩展了之前的分词模型,称为Pitman-Yor Hidden SemiMarkov model (PYHSMM),并被认为是一种直接从字符串中构建类n-gram语言模型的方法,同时集成了字符和词级信息。在日语、中文和泰语的标准数据集上进行的实验结果表明,该方法优于以前的结果,产生了最先进的准确性。该模型还将用于分析非先验词汇的语言结构。
We propose a nonparametric Bayesian model for joint unsupervised word segmentation and part-of-speech tagging from raw strings. Extending a previous model for word segmentation, our model is called a Pitman-Yor Hidden SemiMarkov Model (PYHSMM) and considered as a method to build a class n-gram language model directly from strings, while integrating character and word level information. Experimental results on standard datasets on Japanese, Chinese and Thai revealed it outperforms previous results to yield the state-of-the-art accuracies. This model will also serve to analyze a structure of a language whose words are not identified a priori.