Computational Linguistics and Intelligent Text Processing

Computational Linguistics and Intelligent Text Processing
复制标题

计算语言学与智能文本处理

DOI:
10.1007/978-3-319-75487-1_30
复制
发表时间:
2018
期刊:
--
影响因子:
--
通讯作者:
Simaki V
Simaki V
中科院分区:
--
文献类型:
--
作者:
Simaki V

文献摘要

被引文献

相似文献

在本文中,我们解决了Twitter用户在线文本之后的年龄识别问题。我们使用了一组文本挖掘、基于社会语言学和内容相关的文本特征,并评估了许多众所周知和广泛使用的机器学习分类算法,以检查它们在这项任务中的适用性。实验结果表明,随机森林算法具有优异的性能,准确率达到61%。我们使用ReliefF算法对分类特征的信息量进行排序,并根据年龄语言变异的社会语言学原理对结果进行分析。
In this article, we address the problem of age identification of Twitter users, after their online text. We used a set of text mining, sociolinguistic-based and content-related text features, and we evaluated a number of well-known and widely used machine learning algorithms for classification, in order to examine their appropriateness on this task. The experimental results showed that Random Forest algorithm offered superior performance achieving accuracy equal to 61%. We ranked the classification features after their informativity, using the ReliefF algorithm, and we analyzed the results in terms of the sociolinguistic principles on age linguistic variation.