Predicting Chinese Abbreviations from Definitions: An Empirical Learning Approach Using Support Vector Regression

Predicting Chinese Abbreviations from Definitions: An Empirical Learning Approach Using Support Vector Regression
复制标题

DOI:
10.1007/s11390-008-9156-5
复制
发表时间:
2008-07
影响因子:
0.7
通讯作者:
Xu Sun;Houfeng Wang;Bo Wang-
Xu Sun;Houfeng Wang;Bo Wang-
中科院分区:
--
文献类型:
--
作者:
Xu Sun;Houfeng Wang;Bo Wang-

文献摘要

被引文献

相似文献

在汉语中,短语和命名实体在信息检索中起着核心作用。然而,缩写会降低基于关键字的方法的有效性。本文提出了一种汉语缩略语预测的经验学习方法。在本研究中,将每个缩略语作为相应定义的简化形式(扩展形式),并将缩略语预测形式化为根据相应定义自动生成的缩略语候选之间的评分和排序问题。通过使用支持向量回归(SVR)进行评分,可以得到多个缩略语候选及其SVR值,用于候选排名。实验结果表明,支持向量机方法比目前流行的启发式缩写预测规则具有更好的性能。此外,在缩写预测中,支持向量机方法的性能优于隐马尔可夫模型。
In Chinese, phrases and named entities play a central role in information retrieval. Abbreviations, however, make keyword-based approaches less effective. This paper presents an empirical learning approach to Chinese abbreviation prediction. In this study, each abbreviation is taken as a reduced form of the corresponding definition (expanded form), and the abbreviation prediction is formalized as a scoring and ranking problem among abbreviation candidates, which are automatically generated from the corresponding definition. By employing Support Vector Regression (SVR) for scoring, we can obtain multiple abbreviation candidates together with their SVR values, which are used for candidate ranking. Experimental results show that the SVR method performs better than the popular heuristic rule of abbreviation prediction. In addition, in abbreviation prediction, the SVR method outperforms the hidden Markov model (HMM).