Integrating Generative and Discriminative Character-Based Models for Chinese Word Segmentation

Integrating Generative and Discriminative Character-Based Models for Chinese Word Segmentation
复制标题

集成基于生成和判别字符的模型进行中文分词

DOI:
10.1145/2184436.2184440
复制
发表时间:
2012-06
期刊:
ACM Transactions on Asian Language Information Processing
影响因子:
--
通讯作者:
Keh-Yih Su
Keh-Yih Su
中科院分区:
其他
文献类型:
--
作者:
Kun Wang;Chengqing Zong;Keh-Yih Su

文献摘要

参考文献

被引文献

相似文献

在中文分词的统计方法中,基于词的 n-gram(生成)模型和基于字符的标记(判别)模型是文献中的两种主要方法。前者对于词汇内(IV)单词具有出色的性能;然而,它对词汇外(OOV)单词的处理效果很差。另一方面,尽管后者对于 OOV 词更加稳健,但对于 IV 词却未能提供令人满意的性能。这两种方法由于使用的单位(单词与字符)和采用的模型形式(生成与判别)而表现不同。一般来说,基于字符的方法比基于单词的方法更稳健,因为字符的词汇是一个封闭的集合;判别模型比生成模型更稳健,因为它们可以灵活地包含各种可用信息,例如未来背景。 本文首先提出了一种基于字符的 n-gram 模型来增强生成方法的鲁棒性。然后,所提出的生成模型进一步与基于字符的判别模型集成,以利用这两种方法。我们的实验表明,这种集成方法优于文献中报道的所有现有方法。然后,进行完整而详细的错误分析。由于关键错误的很大一部分与数字/外来字符串有关,因此将字符类型信息合并到模型中以进一步提高其性能。最后,所提出的集成方法在跨域语料库上进行了测试,并提出了半监督域适应算法,并在我们的实验中证明是有效的。
Among statistical approaches to Chinese word segmentation, the word-based n-gram (generative) model and the character-based tagging (discriminative) model are two dominant approaches in the literature. The former gives excellent performance for the in-vocabulary (IV) words; however, it handles out-of-vocabulary (OOV) words poorly. On the other hand, though the latter is more robust for OOV words, it fails to deliver satisfactory performance for IV words. These two approaches behave differently due to the unit they use (word vs. character) and the model form they adopt (generative vs. discriminative). In general, character-based approaches are more robust than word-based ones, as the vocabulary of characters is a closed set; and discriminative models are more robust than generative ones, since they can flexibly include all kinds of available information, such as future context. This article first proposes a character-based n-gram model to enhance the robustness of the generative approach. Then the proposed generative model is further integrated with the character-based discriminative model to take advantage of both approaches. Our experiments show that this integrated approach outperforms all the existing approaches reported in the literature. Afterwards, a complete and detailed error analysis is conducted. Since a significant portion of the critical errors is related to numerical/foreign strings, character-type information is then incorporated into the model to further improve its performance. Last, the proposed integrated approach is tested on cross-domain corpora, and a semi-supervised domain adaptation algorithm is proposed and shown to be effective in our experiments.
DOI: 10.5555/3020488.3020508
发表时间: 2007-07
期刊: --
影响因子: --
作者:
Gholamreza Haffari;Anoop Sarkar
通讯作者: Gholamreza Haffari;Anoop Sarkar
DOI: --
发表时间: 2011-06
期刊: --
影响因子: --
作者:
Weiwei SUN
通讯作者: Weiwei SUN
DOI: 10.3115/1687878.1687951
发表时间: 2009-08
期刊: --
影响因子: --
作者:
Canasai Kruengkrai;Kiyotaka Uchimoto;Jun'ichi Kazama;Yio Wang;Kentaro Torisawa;H. Isahara
通讯作者: Canasai Kruengkrai;Kiyotaka Uchimoto;Jun'ichi Kazama;Yio Wang;Kentaro Torisawa;H. Isahara
DOI: 10.3115/1610075.1610155
发表时间: 2006-07
期刊: --
影响因子: --
作者:
Kristina Toutanova
通讯作者: Kristina Toutanova
DOI: 10.1198/jasa.2008.s236
发表时间: 2008-06
影响因子: 3.7
作者:
T. Burr
通讯作者: T. Burr