Determining Associations with Word Embedding in Heterogeneous Network for Detecting Off-Label Drug Uses

Determining Associations with Word Embedding in Heterogeneous Network for Detecting Off-Label Drug Uses
复制标题

DOI:
10.1109/ichi.2017.78
复制
发表时间:
2017-08
期刊:
2017 IEEE International Conference on Healthcare Informatics (ICHI)
影响因子:
--
通讯作者:
Christopher C. Yang;Mengnan Zhao
Christopher C. Yang;Mengnan Zhao
中科院分区:
其他
文献类型:
--
作者:
Christopher C. Yang;Mengnan Zhao

文献摘要

相似文献

超说明书用药在临床实践中相当普遍,在一定程度上是不可避免的。这些用途可能会提供有效的治疗,并建议临床创新,但由于缺乏科学支持,它们具有未知的风险,可能会导致严重的后果。由于获得有关超说明书药物使用的信息可以为医疗保健专业人员和药物制造商等利益相关者提供线索,以进一步调查药物的疗效和安全性,因此需要开发一种系统的方法来检测超说明书药物使用。考虑到健康消费者之间在线健康社区(OHC)的讨论越来越多,我们建议利用OHC中的大量及时信息来开发一种自动化方法,用于从健康消费者生成的数据中检测标签外药物使用。从文本语料库中,我们提取医疗实体(疾病,药物,药物不良反应)与基于词汇的方法和测量它们的相互作用与词嵌入模型,在此基础上,我们构建了一个异构的医疗保健网络。我们定义了几个基于元路径的指标来描述异构网络中的药物-疾病关联,并将其作为特征来训练基于随机森林算法的二元分类器,以识别已知的药物-疾病关联。分类模型在结合词嵌入特征时获得了更好的结果,并且在使用关联规则挖掘特征和词嵌入特征时获得了最佳性能,F1得分达到0.939,在此基础上,我们确定了2,125种可能的超说明书药物使用,并通过在PubMed和FAERS中搜索证据来检查其潜力。
Off-label drug use is quite common in clinical practice and inevitable to some extent. Such uses might deliver effective treatment and suggest clinical innovation sometimes, however, they have the unknown risk to cause serious outcomes due to lacking scientific support. As gaining information about off-label drug use could present a clue to the stakeholders such as healthcare professionals and medication manufacturers to further the investigation on drug efficacy and safety, it raises the need to develop a systematic way to detect off-label drug uses. Considering the increasing discussions in online health communities (OHCs) among the health consumers, we proposed to harness the large volume of timely information in OHCs to develop an automated method for detecting off-label drug uses from health consumer generated data. From the text corpus, we extracted medical entities (diseases, drugs, and adverse drug reactions) with lexicon-based approaches and measured their interactions with word embedding models, based on which, we constructed a heterogeneous healthcare network. We defined several meta-path-based indicators to describe the drug-disease associations in the heterogeneous network and used them as features to train a binary classifier built on Random Forest algorithm, to recognize the known drug-disease associations. The classification model obtained better results when incorporating word embedding features and achieved the best performance when using both association rule mining features and word embedding features, with F1-score reaching 0.939, based on which, we identified 2,125 possible off-label drug uses and checked their potential by searching evidence in PubMed and FAERS.