Text classification models for the automatic detection of nonmedical prescription medication use from social media.

Text classification models for the automatic detection of nonmedical prescription medication use from social media.
复制标题

DOI:
10.1186/s12911-021-01394-0
复制
发表时间:
2021-01-26
影响因子:
3.5
通讯作者:
Sarker A
Sarker A
中科院分区:
医学3区
文献类型:
--
作者:
Al-Garadi MA;Yang YC;Cai H;Ruan Y;O'Connor K;Graciela GH;Perrone J;Sarker A

文献摘要

参考文献

被引文献

相似文献

处方药(PM)误用/滥用已成为美国的全国性危机,社交媒体已被建议作为进行主动监测的潜在资源。然而,自动化基于社交媒体的监控系统是具有挑战性的,需要先进的自然语言处理(NLP)和机器学习方法。在本文中,我们描述了用于检测Twitter上的PM滥用自我报告的自动文本分类模型的开发和评估。我们试验了最先进的基于双向转换器的语言模型,该模型利用推特级表示来实现迁移学习(例如BERT、RoBERTa、XLNet、AlBERT和DistilBERT),提出了基于融合的方法,并将开发的模型与几种传统机器学习方法(包括深度学习方法)进行了比较。使用公共数据集,我们评估了分类器对非大多数“滥用/误用”类别进行分类的能力。我们提出的基于融合的模型表现明显优于最佳传统模型(f1评分[95% CI]: 0.67 [0.64-0.69] vs. 0.45[0.42-0.48])。我们通过使用不同训练集大小的实验说明,与其他模型相比,基于变压器的模型更稳定,需要更少的注释数据。与过去的方法相比,我们表现最好的分类模型取得了显著的改进,这使得它适合于对Twitter上的非医疗PM使用进行自动连续监控。BERT、类BERT和基于融合的模型优于传统的机器学习和深度学习模型,在过去多年对社交媒体处方药滥用/滥用分类主题的研究中取得了实质性改进,由于呈现非医疗使用信息的独特方式,这已被证明是一项复杂的任务。为了进一步改进BERT和类BERT模型,需要克服与缺乏上下文和社交媒体语言性质相关的几个挑战。这些实验驱动的挑战是未来潜在的研究方向。
Prescription medication (PM) misuse/abuse has emerged as a national crisis in the United States, and social media has been suggested as a potential resource for performing active monitoring. However, automating a social media-based monitoring system is challenging—requiring advanced natural language processing (NLP) and machine learning methods. In this paper, we describe the development and evaluation of automatic text classification models for detecting self-reports of PM abuse from Twitter. We experimented with state-of-the-art bi-directional transformer-based language models, which utilize tweet-level representations that enable transfer learning (e.g., BERT, RoBERTa, XLNet, AlBERT, and DistilBERT), proposed fusion-based approaches, and compared the developed models with several traditional machine learning, including deep learning, approaches. Using a public dataset, we evaluated the performances of the classifiers on their abilities to classify the non-majority “abuse/misuse” class. Our proposed fusion-based model performs significantly better than the best traditional model (F1-score [95% CI]: 0.67 [0.64–0.69] vs. 0.45 [0.42–0.48]). We illustrate, via experimentation using varying training set sizes, that the transformer-based models are more stable and require less annotated data compared to the other models. The significant improvements achieved by our best-performing classification model over past approaches makes it suitable for automated continuous monitoring of nonmedical PM use from Twitter. BERT, BERT-like and fusion-based models outperform traditional machine learning and deep learning models, achieving substantial improvements over many years of past research on the topic of prescription medication misuse/abuse classification from social media, which had been shown to be a complex task due to the unique ways in which information about nonmedical use is presented. Several challenges associated with the lack of context and the nature of social media language need to be overcome to further improve BERT and BERT-like models. These experimental driven challenges are represented as potential future research directions.
药物不良事件的文本挖掘:前景、挑战和最新技术。
DOI: 10.1007/s40264-014-0218-z
发表时间: 2014-10
期刊: DRUG SAFETY
影响因子: 4.2
作者:
Harpaz, Rave;Callahan, Alison;Tamang, Suzanne;Low, Yen;Odgers, David;Finlayson, Sam;Jung, Kenneth;LePendu, Paea;Shah, Nigam H.
通讯作者: Shah, Nigam H.
DOI: 10.2196/jmir.2503
发表时间: 2013-04-17
影响因子: 7.4
作者:
Hanson CL;Burton SH;Giraud-Carrier C;West JH;Barnes MD;Hansen B
通讯作者: Hansen B
DOI: 10.1162/tacl_a_00298
发表时间: 2020-01-01
影响因子: 10.9
作者:
Ettinger, Allyson
通讯作者: Ettinger, Allyson
DOI: 10.3389/fpsyt.2018.00135
发表时间: 2018
影响因子: 4.7
作者:
Chary M;Yi D;Manini AF
通讯作者: Manini AF
DOI: 10.3389/fphar.2018.00791
发表时间: 2018-07-26
影响因子: 5.6
作者:
Bigeard, Elise;Grabar, Natalia;Thiessard, Frantz
通讯作者: Thiessard, Frantz