课题基金 / 基金详情

Study of Class-based Language Model and its Application to Japanese Morphological Analysis

Study of Class-based Language Model and its Application to Japanese Morphological Analysis
基于类的语言模型研究及其在日语词法分析中的应用
批准号:
10680383
负责人:
KITA Kenji
金额:
$1.54万
依托单位:
依托单位国家:
日本
项目类别:
Grant-in-Aid for Scientific Research (C)
财政年份:
1998
资助国家:
日本
项目状态:
已结题
起止时间:
1998 至 1999

项目摘要

项目成果

KITA Kenji的其他基金

相关文献

中文摘要
翻译
词形分析是日语语言处理中最基本的过程。在日语的词法分析中,分词是一个重要的问题,因为它的书写系统中没有标记词的边界。在这个研究项目中,我们首先使用基于字符的n-gram模型研究了一个分词模型,这是我们的基线方法。接下来,我们将PPM*压缩算法应用于分词问题。PPM (Prediction by Partial Matching)是一种基于有限上下文概率建模技术的无损压缩算法,PPM*是PPM的一种变体,其中对上下文长度没有先验限制。然后,我们研究了一种基于字符类模型的分词方法。字符类模型比基于字符的模型更健壮,因为字符类模型的参数数量比基于字符的模型少。日文字符聚类的度量是不同于模型估计的语料库上的熵,搜索方法是基于贪心算法。出于这个原因,这种聚类方法在不给出类别数量的情况下为我们提供了最佳的字符分类。在ADD (ATR Dialogue Database)语料库上的实验结果表明,采用字符类模型的日语分词方法比基于字符的分词方法准确率更高。特别地,所提出的方法采用变长n;-gram类模型对开放文本的查全率为96.38%,查准率为96.23%。
英文摘要
Morphological analysis is the most fundamental process of Japanese language processing. In Japanese morphological analysis, word segmentation is an important problem because word boundaries are not marked in its writing system.In this research project, we first studied a word segmentation model using a character-based n-gram model, which is our baseline method. Next, we applied the PPM* compression algorithm to the problem of word segmentation. PPM (Prediction by Partial Matching) is a lossless compression algorithm based on a finite-context probabilistic modeling technique and PPM* is a variant of PPM, in which there is no a priori bound on context length.We then studied a method for word segmentation based on a character class model. The character class model is more robust than a character-based model because the number of parameters of the character class model is fewer than that of a character-based model. The measurement for Japanese character clustering is the entropy on a corpus different from the corpus for model estimation and the search method is based on the greedy algorithm. For this reason, this clustering method gives us an optimum character classification without giving the number of classes. As the result of experiments on the ADD (ATR Dialogue Database) corpus, the proposed Japanese word segmenter using the character class model marked a higher accuracy than a character-based model. In particular, the proposed method using a variable-length n;-gram class model achieved 96.38% recall and 96.23% precision for open text.
期刊论文(30)
专著(0)
科研奖励(0)
会议论文
Y.Tanaka,K.Kita: "JCKE Multilingual Corpus of Major Asian Languages"Proceedings of TKE'99. 660-670 (1999)
Y.Tanaka,K.Kita:“JCKE 主要亚洲语言多语言语料库”TKE99 论文集。
DOI: --
发表时间:
期刊:
影响因子: --
作者: []
通讯作者:
小田裕樹,北研二: "PPM^*言語モデルを用いた日本語単語分割"情報処理学会論文誌. (印刷中). (2000)
Hiroki Oda、Kenji Kita:“使用 PPM^* 语言模型进行日语分词”,日本信息处理学会会刊(2000 年出版)。
DOI: --
发表时间:
期刊:
影响因子: --
作者: []
通讯作者:
H.Oda,K,Kita: "A Character-Based Japanese Word Segmenter Usirtg PPM^*-Based Langauge Model"Proceedings of ICCPOL'99. 527-532 (1999)
H.Oda,K,Kita:“基于字符的日语分词器 Usirtg PPM^*-基于语言模型”ICCPOL99 论文集。
DOI: --
发表时间:
期刊:
影响因子: --
作者: []
通讯作者:
K.Kita: "Automatic Clustering of Languages Based on Ptobabilistic Models"Journal of Quantitative Linguistics. 6・2. 167-171 (1999)
K.Kita:《基于Ptobabilistic模型的语言自动聚类》定量语言学杂志167-171(1999)。
DOI: --
发表时间:
期刊:
影响因子: --
作者: []
通讯作者:
共 28 条
    Metastasis initiating cell of lung cancer brain metastases
    • 批准号:
      24659403
    • 项目类别:
      Grant-in-Aid for Challenging Exploratory Research
    • 资助金额:
      $2.33万
    • 财政年份:
      2012
    • 负责人:
      KITA Kenji
    • 依托单位:
    Semantic and Affective Multimedia Retrieval using EEG-based Biological Information
    • 批准号:
      21300036
    • 项目类别:
      Grant-in-Aid for Scientific Research (B)
    • 资助金额:
      $10.23万
    • 财政年份:
      2009
    • 负责人:
      KITA Kenji
    • 依托单位:
    Study of Effective Speech Recognition based on the Bi-directional Search Algorithm
    • 批准号:
      07680401
    • 项目类别:
      Grant-in-Aid for Scientific Research (C)
    • 资助金额:
      $1.28万
    • 财政年份:
      1995
    • 负责人:
      KITA Kenji
    • 依托单位: