Paraphrase identification and semantic text similarity analysis in Arabic news tweets using lexical, syntactic, and semantic features

Paraphrase identification and semantic text similarity analysis in Arabic news tweets using lexical, syntactic, and semantic features
复制标题

DOI:
10.1016/j.ipm.2017.01.002
复制
发表时间:
2017-05-01
影响因子:
8.6
通讯作者:
Jararweh, Yaser
Jararweh, Yaser
中科院分区:
计算机科学1区
文献类型:
--
作者:
Al-Smadi, Mohammad;Jaradat, Zain;Jararweh, Yaser

文献摘要

被引文献

相似文献

数字信息的快速增长提出了相当大的挑战,特别是在自动化内容分析方面。 Twitter 等社交媒体分享了大量用户的事件、观点、个性等信息。释义识别 (PI) 关注的是识别两个文本是否具有相同/相似的含义,而语义文本相似度 (STS) 关注的是相似程度。这项研究提出了一种最先进的方法,用于阿拉伯语新闻推文中的释义识别和语义文本相似性分析。该方法采用文本处理、特征提取和文本分类的几个阶段。提取词汇、句法和语义特征,以克服当前技术在解决阿拉伯语这些任务时的弱点和限制。最大熵 (MaxEnt) 和支持向量回归 (SVR) 分类器使用这些特征进行训练,并使用为本研究准备的数据集进行评估。实验结果表明,与基线结果相比,该方法取得了良好的结果。 (c) 2017 Elsevier Ltd. 保留所有权利。
The rapid growth in digital information has raised considerable challenges in particular when it comes to automated content analysis. Social media such as twitter share a lot of its users' information about their events, opinions, personalities, etc. Paraphrase Identification (PI) is concerned with recognizing whether two texts have the same/similar meaning, whereas the Semantic Text Similarity (STS) is concerned with the degree of that similarity. This research proposes a state-of-the-art approach for paraphrase identification and semantic text similarity analysis in Arabic news tweets. The approach adopts several phases of text processing, features extraction and text classification. Lexical, syntactic, and semantic features are extracted to overcome the weakness and limitations of the current technologies in solving these tasks for the Arabic language. Maximum Entropy (MaxEnt) and Support Vector Regression (SVR) classifiers are trained using these features and are evaluated using a dataset prepared for this research. The experimentation results show that the approach achieves good results in comparison to the baseline results. (c) 2017 Elsevier Ltd. All rights reserved.