Unknown Word Detection for Chinese by a Corpus-based Learning Method

Unknown Word Detection for Chinese by a Corpus-based Learning Method
复制标题

DOI:
--
复制
发表时间:
1998
期刊:
Int. J. Comput. Linguistics Chin. Lang. Process.
影响因子:
--
通讯作者:
Keh-Jiann Chen;Ming-Hong Bai
Keh-Jiann Chen;Ming-Hong Bai
中科院分区:
其他
文献类型:
--
作者:
Keh-Jiann Chen;Ming-Hong Bai

文献摘要

被引文献

相似文献

汉语计算机处理中最突出的问题之一是句子中的词的识别。由于没有空白来标记单词边界,因此由于分割歧义和词汇表外单词的出现(即,未知的字)。在本文中,提出了一种基于语料库的学习方法,推导出一套句法规则,适用于区分单音节词的单音节语素,这可能是未知的单词或印刷错误的一部分。基于语料库的学习方法具有以下优点:1。自动规则学习,2.自动评估每个规则的性能,以及3.通过动态规则集选择来平衡查全率和查准率。实验结果表明,使用该方法得到的规则集优于手工制作的规则,由人类专家在检测未登录词。
One of the most prominent problems in computer processing of the Chinese language is identification of the words in a sentence. Since there are no blanks to mark word boundaries, identifying words is difficult because of segmentation ambiguities and occurrences of out-of-vocabulary words (i.e., unknown words). In this paper, a corpus-based learning method is proposed which derives sets of syntactic rules that are applied to distinguish monosyllabic words from monosyllabic morphemes which may be parts of unknown words or typographical errors. The corpus-based learning approach has the advantages of: 1. automatic rule learning, 2. automatic evaluation of the performance of each rule, and 3. balancing of recall and precision rates through dynamic rule set selection. The experimental results show that the rule set derived using the proposed method outperformed hand-crafted rules produced by human experts in detecting unknown words.