Pharmacovigilance from social media: mining adverse drug reaction mentions using sequence labeling with word embedding cluster features.

Pharmacovigilance from social media: mining adverse drug reaction mentions using sequence labeling with word embedding cluster features.
复制标题

DOI:
10.1093/jamia/ocu041
复制
发表时间:
2015-05
期刊:
Journal of the American Medical Informatics Association : JAMIA
影响因子:
--
通讯作者:
Gonzalez G
Gonzalez G
中科院分区:
其他
文献类型:
--
作者:
Nikfarjam A;Sarker A;O'Connor K;Ginn R;Gonzalez G

文献摘要

参考文献

被引文献

相似文献

作为分享个人健康信息的平台,社交媒体正变得越来越受欢迎。通过使用自然语言处理(NLP)技术,这些信息可用于公共卫生监测任务,特别是用于药物警戒。然而,社交媒体中的语言是非正式的,用户表达的医学概念通常是非技术性的,描述性的,很难提取。在应对这些挑战方面进展有限,到目前为止,先进的基于机器学习的自然语言处理技术尚未得到充分利用。我们的目标是设计一种基于机器学习的方法,从社交媒体中高度非正式的文本中提取药物不良反应(adr)。方法介绍一种基于机器学习的概念提取系统ADRMine,该系统使用条件随机场(CRFs)。ADRMine利用了多种功能,包括一个用于建模单词语义相似度的新功能。相似性是通过使用深度学习技术,基于无监督、预训练的词表示向量(嵌入)聚类词来建模的,这些词来自社交媒体中未标记的用户帖子。结果ADRMine在ADR提取任务中优于几个强基线系统,f值为0.82。特征分析表明,提出的聚类特征显著提高了提取性能。结论从非正式的、用户生成的内容中提取复杂的医学概念是可能的,并且性能相对较高。我们的方法特别具有可扩展性,适合于社交媒体挖掘,因为它依赖于大量未标记的数据,从而减少了对大型、带注释的训练数据集的需求。
Objective Social media is becoming increasingly popular as a platform for sharing personal health-related information. This information can be utilized for public health monitoring tasks, particularly for pharmacovigilance, via the use of natural language processing (NLP) techniques. However, the language in social media is highly informal, and user-expressed medical concepts are often nontechnical, descriptive, and challenging to extract. There has been limited progress in addressing these challenges, and thus far, advanced machine learning-based NLP techniques have been underutilized. Our objective is to design a machine learning-based approach to extract mentions of adverse drug reactions (ADRs) from highly informal text in social media. Methods We introduce ADRMine, a machine learning-based concept extraction system that uses conditional random fields (CRFs). ADRMine utilizes a variety of features, including a novel feature for modeling words’ semantic similarities. The similarities are modeled by clustering words based on unsupervised, pretrained word representation vectors (embeddings) generated from unlabeled user posts in social media using a deep learning technique. Results ADRMine outperforms several strong baseline systems in the ADR extraction task by achieving an F-measure of 0.82. Feature analysis demonstrates that the proposed word cluster features significantly improve extraction performance. Conclusion It is possible to extract complex medical concepts, with relatively high performance, from informal, user-generated content. Our approach is particularly scalable, suitable for social media mining, as it relies on large volumes of unlabeled data, thus diminishing the need for large, annotated training data sets.
DOI: 10.1016/j.jbi.2011.07.005
发表时间: 2011-12
影响因子: 4.5
作者:
Benton, Adrian;Ungar, Lyle;Hill, Shawndra;Hennessy, Sean;Mao, Jun;Chung, Annie;Leonard, Charles E.;Holmes, John H.
通讯作者: Holmes, John H.
DOI: 10.1038/msb.2009.98
发表时间: 2010
影响因子: 9.9
作者:
Kuhn M;Campillos M;Letunic I;Jensen LJ;Bork P
通讯作者: Bork P
DOI: 10.1093/database/bau084
发表时间: 2014-09-10
影响因子: 5.8
作者:
Emadzadeh, Ehsan;Nikfarjam, Azadeh;Gonzalez, Graciela
通讯作者: Gonzalez, Graciela
DOI: 10.1186/2041-1480-3-15
发表时间: 2012-12-20
影响因子: 1.9
作者:
Gurulingappa H;Mateen-Rajput A;Toldo L
通讯作者: Toldo L
DOI: 10.1177/001316446002000104
发表时间: 1960-01-01
影响因子: 2.7
作者:
COHEN, J
通讯作者: COHEN, J