Validating Machine Learning Algorithms for Twitter Data Against Established Measures of Suicidality.

Validating Machine Learning Algorithms for Twitter Data Against Established Measures of Suicidality.
复制标题

DOI:
10.2196/mental.4822
复制
发表时间:
2016-05-16
期刊:
影响因子:
5.2
通讯作者:
Hanson CL
Hanson CL
中科院分区:
医学2区
文献类型:
--
作者:
Braithwaite SR;Giraud-Carrier C;West J;Barnes MD;Hanson CL

文献摘要

被引文献

相似文献

在美国,自杀是导致死亡的主要原因之一,需要新的评估方法来真实的跟踪其风险。我们的目标是验证机器学习算法在Twitter数据中的使用,以验证美国人口中自杀倾向的经验验证措施。使用机器学习算法,将135名Mechanical Turk(MTurk)参与者的Twitter提要与经过验证的自杀风险自我报告措施进行了比较。我们的研究结果表明,通过机器学习算法可以很容易地将自杀风险高的人与那些没有自杀风险的人区分开来,这些算法在92%的病例中准确地识别出具有临床意义的自杀率(灵敏度:53%,特异性:97%,阳性预测值:75%,阴性预测值:93%)。机器学习算法可以有效地区分有自杀风险的人和没有自杀风险的人。自杀倾向的证据可以在非临床人群中使用社交媒体数据进行测量。
One of the leading causes of death in the United States (US) is suicide and new methods of assessment are needed to track its risk in real time. Our objective is to validate the use of machine learning algorithms for Twitter data against empirically validated measures of suicidality in the US population. Using a machine learning algorithm, the Twitter feeds of 135 Mechanical Turk (MTurk) participants were compared with validated, self-report measures of suicide risk. Our findings show that people who are at high suicidal risk can be easily differentiated from those who are not by machine learning algorithms, which accurately identify the clinically significant suicidal rate in 92% of cases (sensitivity: 53%, specificity: 97%, positive predictive value: 75%, negative predictive value: 93%). Machine learning algorithms are efficient in differentiating people who are at a suicidal risk from those who are not. Evidence for suicidality can be measured in nonclinical populations using social media data.