Inducing Word and Part-of-Speech with Pitman-Yor Hidden Semi-Markov Models
Inducing Word and Part-of-Speech with Pitman-Yor Hidden Semi-Markov Models
复制标题
DOI:
10.3115/v1/p15-1171
复制
发表时间:
2015-07
期刊:
影响因子:
--
通讯作者:
Kei Uchiumi;Hiroshi Tsukahara;D. Mochihashi
中科院分区:
文献类型:
--
作者:
Kei Uchiumi;Hiroshi Tsukahara;D. Mochihashi
We propose a nonparametric Bayesian model for joint unsupervised word segmentation and part-of-speech tagging from raw strings. Extending a previous model for word segmentation, our model is called a Pitman-Yor Hidden SemiMarkov Model (PYHSMM) and considered as a method to build a class n-gram language model directly from strings, while integrating character and word level information. Experimental results on standard datasets on Japanese, Chinese and Thai revealed it outperforms previous results to yield the state-of-the-art accuracies. This model will also serve to analyze a structure of a language whose words are not identified a priori.