Named Entity Recognition in Vietnamese documents

Named Entity Recognition in Vietnamese documents
复制标题

DOI:
10.2201/niipi.2007.4.2
复制
发表时间:
2007-03
期刊:
Progress in Informatics
影响因子:
--
通讯作者:
Q. Tran;T. Pham;Quoc Hung Ngo;D. Dinh;Nigel Collier
Q. Tran;T. Pham;Quoc Hung Ngo;D. Dinh;Nigel Collier
中科院分区:
其他
文献类型:
--
作者:
Q. Tran;T. Pham;Quoc Hung Ngo;D. Dinh;Nigel Collier

文献摘要

被引文献

相似文献

命名实体识别 (NER) 旨在将文档中的单词分类为预定义的目标实体类,现在被认为是许多自然语言处理任务的基础,例如信息检索、机器翻译、信息提取和问答。本文介绍了将基于支持向量机 (SVM) 的 NER 模型应用于越南语的实验结果。尽管这种最先进的机器学习方法已广泛应用于几种经过深入研究的语言的 NER 中,但这是该方法首次应用于越南语。在与条件随机场 (CRF) 的比较中,SVM 模型通过优化其特征窗口大小,表现优于 CRF,总体 F 得分为 87.75。本文还详细讨论了越南语的特征,并分析了影响该任务表现的因素。
NamedEntityRecognition (NER) aimstoclassify wordsin a documentintopre-definedtarget entity classes and is now considered to be fundamental for many natural language processing tasks such a si nformation retrieval, machine translation, information extraction and question answering. This paper presents the results of an experiment in which a Support Vector Machine (SVM) based NER model is applied to the Vietnamese language. Though this state of the art machine learning method has been widely applied to NER in several well-studied languages, this is the first time this method has been applied to Vietnamese. In a comparison against Conditional Random Fields (CRFs) the SVM model was shown to outperform CRF by optimizing its feature window size, obtaining an overall F-score of 87.75. The paper also presents a detailed discussion about the characteristics of the Vietnamese language and provides an analysis of the factors which influence performance in this task.