Chinese Unknown Word Identification Using Character-based Tagging and Chunking

Chinese Unknown Word Identification Using Character-based Tagging and Chunking
复制标题

DOI:
10.3115/1075178.1075215
复制
发表时间:
2003-07
期刊:
--
影响因子:
--
通讯作者:
Chooi-Ling Goh;Masayuki Asahara;Yuji Matsumoto
Chooi-Ling Goh;Masayuki Asahara;Yuji Matsumoto
中科院分区:
其他
文献类型:
--
作者:
Chooi-Ling Goh;Masayuki Asahara;Yuji Matsumoto

文献摘要

被引文献

相似文献

由于书面汉语没有空间来划分单词,因此切分中文文本成为一项必不可少的任务。在这个任务中,出现了未知词的问题。不可能在字典中注册所有单词,因为新单词总是可以通过组合字符来创建。我们提出了一个统一的解决方案来检测中文文本中的未登录词。首先,进行形态分析,以获得初始分割和POS标签,然后使用分块器检测未登录词。
Since written Chinese has no space to delimit words, segmenting Chinese texts becomes an essential task. During this task, the problem of unknown word occurs. It is impossible to register all words in a dictionary as new words can always be created by combining characters. We propose a unified solution to detect unknown words in Chinese texts. First, a morphological analysis is done to obtain initial segmentation and POS tags and then a chunker is used to detect unknown words.