ReportAGE: Automatically extracting the exact age of Twitter users based on self-reports in tweets.

ReportAGE: Automatically extracting the exact age of Twitter users based on self-reports in tweets.
复制标题

DOI:
10.1371/journal.pone.0262087
复制
发表时间:
2022
期刊:
影响因子:
3.7
通讯作者:
Gonzalez-Hernandez G
Gonzalez-Hernandez G
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Klein AZ;Magge A;Gonzalez-Hernandez G

文献摘要

参考文献

被引文献

相似文献

提高社交媒体数据在研究应用中的效用需要自动检测社交媒体研究人群的人口统计信息(包括用户年龄)的方法。这项研究的目的是开发和评估一种根据推文中的自我报告自动识别用户确切年龄的方法。我们的端到端自动自然语言处理 (NLP) 管道 ReportAGE 包括用于检索可能提及年龄的推文的查询模式、用于区分检索到的自动报告用户确切年龄的推文(“年龄”推文)和不报告用户确切年龄的推文(“无年龄”推文)的分类器,以及用于识别年龄的基于规则的提取。为了开发和评估 ReportAGE,我们手动注释了 11,000 条与查询模式匹配的推文。基于所有五位注释者注释的 1000 条推文,区分“年龄”和“无年龄”推文的注释者间一致性(Fleiss’ kappa)为 0.80,识别注释者一致同意的“年龄”推文中的确切年龄为 0.95。基于 RoBERTa-Large 预训练 Transformer 模型的深度神经网络分类器在“年龄”类别中获得了最高的 F1 分数 0.914(精度 = 0.905,召回率 = 0.942)。当使用分类器的预测评估年龄提取时,“年龄”类别的 F1 分数为 0.855(精度 = 0.805,召回率 = 0.914)。当直接在保留的测试集上对其进行评估时,“年龄”类别的 F1 分数为 0.931(精度 = 0.873,召回率 = 0.998)。我们在 245,927 位用户发布的超过 12 亿条推文集合上部署了 ReportAGE,并预测了其中 132,637 名用户 (54%) 的年龄。将精确年龄的检测扩展到如此大量的用户可以提高社交媒体数据在与现有二元或多类分类方法的预定义年龄分组不一致的研究应用中的效用。
Advancing the utility of social media data for research applications requires methods for automatically detecting demographic information about social media study populations, including users’ age. The objective of this study was to develop and evaluate a method that automatically identifies the exact age of users based on self-reports in their tweets. Our end-to-end automatic natural language processing (NLP) pipeline, ReportAGE, includes query patterns to retrieve tweets that potentially mention an age, a classifier to distinguish retrieved tweets that self-report the user’s exact age (“age” tweets) and those that do not (“no age” tweets), and rule-based extraction to identify the age. To develop and evaluate ReportAGE, we manually annotated 11,000 tweets that matched the query patterns. Based on 1000 tweets that were annotated by all five annotators, inter-annotator agreement (Fleiss’ kappa) was 0.80 for distinguishing “age” and “no age” tweets, and 0.95 for identifying the exact age among the “age” tweets on which the annotators agreed. A deep neural network classifier, based on a RoBERTa-Large pretrained transformer model, achieved the highest F1-score of 0.914 (precision = 0.905, recall = 0.942) for the “age” class. When the age extraction was evaluated using the classifier’s predictions, it achieved an F1-score of 0.855 (precision = 0.805, recall = 0.914) for the “age” class. When it was evaluated directly on the held-out test set, it achieved an F1-score of 0.931 (precision = 0.873, recall = 0.998) for the “age” class. We deployed ReportAGE on a collection of more than 1.2 billion tweets, posted by 245,927 users, and predicted ages for 132,637 (54%) of them. Scaling the detection of exact age to this large number of users can advance the utility of social media data for research applications that do not align with the predefined age groupings of extant binary or multi-class classification approaches.
DOI: 10.1371/journal.pone.0115545
发表时间: 2015
期刊: PloS one
影响因子: 3.7
作者:
Sloan L;Morgan J;Burnap P;Williams M
通讯作者: Williams M
根据语言和元数据功能预测Twitter用户的年龄组。
DOI: 10.1371/journal.pone.0183537
发表时间: 2017
期刊: PloS one
影响因子: 3.7
作者:
Morgan-Lopez AA;Kim AE;Chew RF;Ruddle P
通讯作者: Ruddle P
DOI: 10.1016/j.yjbinx.2020.100076
发表时间: 2020-01-01
影响因子: 4.5
作者:
Klein, Ari Z;Cai, Haitao;Gonzalez-Hernandez, Graciela
通讯作者: Gonzalez-Hernandez, Graciela
DOI: 10.1108/00330330610681286
发表时间: 2006-01-01
影响因子: --
作者:
Porter, M. F.
通讯作者: Porter, M. F.
DOI: 10.1007/s40264-018-0731-6
发表时间: 2019-03
期刊: Drug safety
影响因子: 4.2
作者:
Golder S;Chiuve S;Weissenbacher D;Klein A;O'Connor K;Bland M;Malin M;Bhattacharya M;Scarazzini LJ;Gonzalez-Hernandez G
通讯作者: Gonzalez-Hernandez G