Study of Class-based Language Model and its Application to Japanese Morphological Analysis
Study of Class-based Language Model and its Application to Japanese Morphological Analysis
批准号:
10680383
负责人:
KITA Kenji
金额:
$1.54万
依托单位国家:
日本
项目类别:
Grant-in-Aid for Scientific Research (C)
财政年份:
1998
资助国家:
日本
项目状态:
已结题
起止时间:
1998 至 1999
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Morphological analysis is the most fundamental process of Japanese language processing. In Japanese morphological analysis, word segmentation is an important problem because word boundaries are not marked in its writing system.In this research project, we first studied a word segmentation model using a character-based n-gram model, which is our baseline method. Next, we applied the PPM* compression algorithm to the problem of word segmentation. PPM (Prediction by Partial Matching) is a lossless compression algorithm based on a finite-context probabilistic modeling technique and PPM* is a variant of PPM, in which there is no a priori bound on context length.We then studied a method for word segmentation based on a character class model. The character class model is more robust than a character-based model because the number of parameters of the character class model is fewer than that of a character-based model. The measurement for Japanese character clustering is the entropy on a corpus different from the corpus for model estimation and the search method is based on the greedy algorithm. For this reason, this clustering method gives us an optimum character classification without giving the number of classes. As the result of experiments on the ADD (ATR Dialogue Database) corpus, the proposed Japanese word segmenter using the character class model marked a higher accuracy than a character-based model. In particular, the proposed method using a variable-length n;-gram class model achieved 96.38% recall and 96.23% precision for open text.
期刊论文(30)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Y.Tanaka,K.Kita: "JCKE Multilingual Corpus of Major Asian Languages"Proceedings of TKE'99. 660-670 (1999)
Y.Tanaka,K.Kita:“JCKE 主要亚洲语言多语言语料库”TKE99 论文集。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
小田裕樹,北研二: "PPM^*言語モデルを用いた日本語単語分割"情報処理学会論文誌. (印刷中). (2000)
Hiroki Oda、Kenji Kita:“使用 PPM^* 语言模型进行日语分词”,日本信息处理学会会刊(2000 年出版)。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
H.Oda,K,Kita: "A Character-Based Japanese Word Segmenter Usirtg PPM^*-Based Langauge Model"Proceedings of ICCPOL'99. 527-532 (1999)
H.Oda,K,Kita:“基于字符的日语分词器 Usirtg PPM^*-基于语言模型”ICCPOL99 论文集。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
K.Kita: "Automatic Clustering of Languages Based on Ptobabilistic Models"Journal of Quantitative Linguistics. 6・2. 167-171 (1999)
K.Kita:《基于Ptobabilistic模型的语言自动聚类》定量语言学杂志167-171(1999)。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
北 研二: "確率的言語モデル"東京大学出版会. 256 (1999)
Kenji Kita:《概率语言模型》东京大学出版社 256 (1999)。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
共 28 条
Metastasis initiating cell of lung cancer brain metastases
-
批准号:24659403
-
项目类别:Grant-in-Aid for Challenging Exploratory Research
-
资助金额:$2.33万
-
财政年份:2012
-
负责人:KITA Kenji
-
依托单位:
Semantic and Affective Multimedia Retrieval using EEG-based Biological Information
-
批准号:21300036
-
项目类别:Grant-in-Aid for Scientific Research (B)
-
资助金额:$10.23万
-
财政年份:2009
-
负责人:KITA Kenji
-
依托单位:
Study of Effective Speech Recognition based on the Bi-directional Search Algorithm
-
批准号:07680401
-
项目类别:Grant-in-Aid for Scientific Research (C)
-
资助金额:$1.28万
-
财政年份:1995
-
负责人:KITA Kenji
-
依托单位: