Predicting age groups of Twitter users based on language and metadata features.

Predicting age groups of Twitter users based on language and metadata features.
复制标题

根据语言和元数据功能预测Twitter用户的年龄组。

DOI:
10.1371/journal.pone.0183537
复制
发表时间:
2017
期刊:
影响因子:
3.7
通讯作者:
Ruddle P
Ruddle P
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Morgan-Lopez AA;Kim AE;Chew RF;Ruddle P

文献摘要

参考文献

被引文献

相似文献

卫生组织越来越多地使用 Twitter 等社交媒体向目标受众传播健康信息。确定目标受众(例如年龄组)的覆盖程度对于评估社交媒体教育活动的影响至关重要。本研究的主要目的是检验语言和元数据特征在预测 Twitter 用户年龄方面的单独和联合预测有效性。我们通过使用 Twitter 搜索应用程序编程接口收集公开的生日公告推文,创建了不同年龄段(青少年、年轻人、成人)的 Twitter 用户的标记数据集。我们手动审核结果,并针对每个带有年龄标签的用户名,收集了 200 条最新公开的推文和用户用户名的元数据。标记数据被分为训练数据集和测试数据集。我们创建了单独的模型来检查仅语言特征、仅元数据特征、语言和元数据特征以及来自另一个经过年龄验证的数据集的单词/短语的预测有效性。我们估计了每个模型的准确率、精确率、召回率和 F1 指标。对每个年龄组进行 L1 正则化逻辑回归模型,并比较每个年龄组的训练集和测试集之间的预测概率。计算 Cohen 的 d 效应大小以检查显着特征的相对重要性。同时包含推文语言特征和元数据特征的模型表现最好(准确率 74%、召回率 74%、F1 74%),而仅包含 Twitter 元数据特征的模型准确度最差(准确率 58%、召回率 60%、F1 分数 57%)。最重要的预测特征包括使用诸如针对青少年的“学校”和针对年轻人的“学院”等术语。总体而言,准确预测老年人更具挑战性。这些结果表明,检查语言和 Twitter 元数据特征来预测青少年和年轻成人 Twitter 用户可能有助于为公共卫生监测和评估研究提供信息。
Health organizations are increasingly using social media, such as Twitter, to disseminate health messages to target audiences. Determining the extent to which the target audience (e.g., age groups) was reached is critical to evaluating the impact of social media education campaigns. The main objective of this study was to examine the separate and joint predictive validity of linguistic and metadata features in predicting the age of Twitter users. We created a labeled dataset of Twitter users across different age groups (youth, young adults, adults) by collecting publicly available birthday announcement tweets using the Twitter Search application programming interface. We manually reviewed results and, for each age-labeled handle, collected the 200 most recent publicly available tweets and user handles’ metadata. The labeled data were split into training and test datasets. We created separate models to examine the predictive validity of language features only, metadata features only, language and metadata features, and words/phrases from another age-validated dataset. We estimated accuracy, precision, recall, and F1 metrics for each model. An L1-regularized logistic regression model was conducted for each age group, and predicted probabilities between the training and test sets were compared for each age group. Cohen’s d effect sizes were calculated to examine the relative importance of significant features. Models containing both Tweet language features and metadata features performed the best (74% precision, 74% recall, 74% F1) while the model containing only Twitter metadata features were least accurate (58% precision, 60% recall, and 57% F1 score). Top predictive features included use of terms such as “school” for youth and “college” for young adults. Overall, it was more challenging to predict older adults accurately. These results suggest that examining linguistic and Twitter metadata features to predict youth and young adult Twitter users may be helpful for informing public health surveillance and evaluation research.
DOI: 10.2196/jmir.4466
发表时间: 2015-11-06
影响因子: 7.4
作者:
Kim AE;Hopper T;Simpson S;Nonnemaker J;Lieberman AJ;Hansen H;Guillory J;Porter L
通讯作者: Porter L
DOI: 10.1037/a0035048
发表时间: 2014-01-01
影响因子: 4
作者:
Kern, Margaret L.;Eichstaedt, Johannes C.;Seligman, Martin E. P.
通讯作者: Seligman, Martin E. P.
DOI: 10.1371/journal.pone.0115545
发表时间: 2015
期刊: PloS one
影响因子: 3.7
作者:
Sloan L;Morgan J;Burnap P;Williams M
通讯作者: Williams M
DOI: 10.15288/jsad.2016.77.349
发表时间: 2016-03-01
影响因子: 3.4
作者:
Cabrera-Nguyen, E. Peter;Cavazos-Rehg, Patricia;Moreno, Megan A.
通讯作者: Moreno, Megan A.
DOI: 10.1371/journal.pone.0073791
发表时间: 2013
期刊: PloS one
影响因子: 3.7
作者:
Schwartz HA;Eichstaedt JC;Kern ML;Dziurzynski L;Ramones SM;Agrawal M;Shah A;Kosinski M;Stillwell D;Seligman ME;Ungar LH
通讯作者: Ungar LH