A Robust Morpheme Sequence and Convolutional Neural Network-Based Uyghur and Kazakh Short Text Classification

A Robust Morpheme Sequence and Convolutional Neural Network-Based Uyghur and Kazakh Short Text Classification
复制标题

基于鲁棒语素序列和卷积神经网络的维吾尔语和哈萨克语短文本分类

DOI:
10.3390/info10120387
复制
发表时间:
2019
影响因子:
--
通讯作者:
Askar Hamdulla
Askar Hamdulla
中科院分区:
--
文献类型:
--
作者:
Sardar Parhat;Mijit Ablimit;Askar Hamdulla

文献摘要

被引文献

相似文献

本文基于多语种词法分析器,对相似度低的维吾尔语和哈萨克语短文本分类进行了研究。一般来说,这些语言的在线语言资源是嘈杂的。因此,预处理是必要的,可以显着提高精度。维吾尔语和哈萨克语是派生形态语言,词是由词干加后缀构成的。在这些语言中,术语通常被用作文本内容的表示,而不包括作为停止词的功能部分。通过提取词干,我们可以收集必要的术语并排除停用词。词素分割工具可以将文本分割成词素,具有95%的高可靠性。在准备了基于单词和词素的训练文本语料库之后,我们应用卷积神经网络(CNN)作为特征选择和文本分类算法来执行文本分类任务。实验结果表明,基于词素的方法优于基于词的方法。词嵌入技术经常用于神经网络框架下的文本表示和作为值表达式,可以将语言单元映射到基于上下文的序列向量空间中,是一种提取和预测词汇表外(OOV)的自然方法。从上下文信息。多语种词法分析为维吾尔语、哈萨克语等低资源语言的处理提供了一种方便的途径。
In this paper, based on the multilingual morphological analyzer, we researched the similar low-resource languages, Uyghur and Kazakh, short text classification. Generally, the online linguistic resources of these languages are noisy. So a preprocessing is necessary and can significantly improve the accuracy. Uyghur and Kazakh are the languages with derivational morphology, in which words are coined by stems concatenated with suffixes. Usually, terms are used as the representation of text content while excluding functional parts as stop words in these languages. By extracting stems we can collect necessary terms and exclude stop words. Morpheme segmentation tool can split text into morphemes with 95% high reliability. After preparing both word- and morpheme-based training text corpora, we apply convolutional neural network (CNN) as a feature selection and text classification algorithm to perform text classification tasks. Experimental results show that the morpheme-based approach outperformed the word-based approach. Word embedding technique is frequently used in text representation both in the framework of neural networks and as a value expression, and can map language units into a sequential vector space based on context, and it is a natural way to extract and predict out-of-vocabulary (OOV) from context information. Multilingual morphological analysis has provided a convenient way for processing tasks of low resource languages like Uyghur and Kazakh.