Computational Linguistics and Intelligent Text Processing
Computational Linguistics and Intelligent Text Processing
复制标题
计算语言学与智能文本处理
DOI:
10.1007/978-3-319-75487-1_30
复制
发表时间:
2018
期刊:
影响因子:
--
通讯作者:
Simaki V
中科院分区:
文献类型:
--
作者:
Simaki V
In this article, we address the problem of age identification of Twitter users, after their online text. We used a set of text mining, sociolinguistic-based and content-related text features, and we evaluated a number of well-known and widely used machine learning algorithms for classification, in order to examine their appropriateness on this task. The experimental results showed that Random Forest algorithm offered superior performance achieving accuracy equal to 61%. We ranked the classification features after their informativity, using the ReliefF algorithm, and we analyzed the results in terms of the sociolinguistic principles on age linguistic variation.