A HowNet-Based Semantic Relatedness Kernel for Text Classification

A HowNet-Based Semantic Relatedness Kernel for Text Classification
复制标题

DOI:
10.11591/telkomnika.v11i4.2361
复制
发表时间:
2013-04
影响因子:
--
通讯作者:
P. Zhang
P. Zhang
中科院分区:
--
文献类型:
--
作者:
P. Zhang

文献摘要

被引文献

相似文献

在文本检索和信息管理中,语义相关核的开发一直是一个很有吸引力的课题。通常,在文本分类中,使用词袋(BOW)方法在向量空间中表示文档。BOW方法不考虑语义相关性信息。为了进一步提高文本分类性能,提出了一种新的基于语义核的支持向量机文本分类算法。该方法首先利用CHI方法选择文档特征向量,然后利用TF-IDF方法计算特征向量的权重,最后利用包含语义相似度计算和语义相关度计算的语义相关度核对文档进行支持向量机分类。实验结果表明,与传统的支持向量机算法相比,该算法在文本分类中实现了改进的分类F1-测度。DOI:http://dx.doi.org/10.11591/telkomnika.v11i4.2361网站
The exploitation of the semantic relatedness kernel has always been an appealing subject in the context of text retrieval and information management. Typically, in text classification the documents are represented in the vector space using the bag-of-words (BOW) approach. The BOW approach does not take into account the semantic relatedness information. To further improve the text classification performance, this paper presents a new semantic-based kernel of support vector machine algorithm for text classification. This method firstly using CHI method to select document feature vectors, secondly calculates the feature vector weights using TF-IDF method, and utilizes the semantic relatedness kernel which involves the semantic similarity computation and semantic relevance computation to classify the document using support vector machines. Experimental results show that compared with the traditional support vector machine algorithm, the algorithm in the text classification achieves improved classification F1-measure. DOI: http://dx.doi.org/10.11591/telkomnika.v11i4.2361