Explainable detection of adverse drug reaction with imbalanced data distribution.

Explainable detection of adverse drug reaction with imbalanced data distribution.
复制标题

DOI:
10.1371/journal.pcbi.1010144
复制
发表时间:
2022-06
影响因子:
4.3
通讯作者:
--
中科院分区:
生物学2区
文献类型:
--
作者:

文献摘要

参考文献

相似文献

对健康相关文本的分析可用于检测药物不良反应(ADR)。ADR检测的最大挑战在于不平衡的数据分布,其中与ADR症状相关的词通常是少数类别。因此,经过训练的模型往往会收敛到一个点,这个点会强烈偏向多数类,然后忽略少数类。由于最常用的交叉熵标准是对准确性的近似,因此该模型更容易集中在多数类上以实现高准确性。为了解决这个问题,现有的方法应用过采样或下采样策略来平衡数据分布并利用少数类中最困难的样本。然而,在序列标记任务中单独增加或减少单个标记的数量将导致句子的句法关系的丢失。针对数据不平衡序列标注任务,提出了一种条件随机场(CRF)的加权变体。这种加权策略可以缓解多数和少数类别之间的数据分布不平衡。CRF可以捕获标记之间的标签关系,而不是在输出层使用softmax。局部可解释的模型不可知解释(LIME)算法被应用于研究模型与加权损失函数之间的性能差异。两个不同ADR任务的实验结果表明,该模型优于先前提出的序列标记方法。上市后药物安全性监测提供了检测导致住院的严重ADR和患者发生的ADR的机会,例如,高并发症患者和接受仅在医院给予的药物的患者。这种监测传统上是通过调查用户来完成的。最近,社交媒体上自动记录用户的ADR可以极大地帮助生物制药企业改进产品。以往自然语言处理中的名称实体识别方法通常是在数据分布均衡的语料库上进行的。相反,ADR检测的数据集非常不平衡。因此,检测器往往会忽略更重要的ADR症状和相关适应症。在这项研究中,我们提出了一个加权CRF模型的基础上BERT的检测任务的ADR。维特比算法的加权变体被实现为向少数类分配更多权重,迫使模型更多地关注少数类以确保有效检测。结果表明,所提出的方法提供了一个显着的性能提升,而不改变不平衡数据任务的模型架构。
Analysis of health-related texts can be used to detect adverse drug reactions (ADR). The greatest challenge for ADR detection lies in imbalanced data distributions where words related to ADR symptoms are often minority classes. As a result, trained models tend to converge to a point that strongly biases towards the majority class and then ignores the minority class. Since the most used cross-entropy criteria is an approximation to accuracy, the model focuses more readily on the majority class to achieve high accuracy. To address this issue, existing methods apply either oversampling or down-sampling strategies to balance the data distribution and exploit the most difficult samples of the minority class. However, increasing or reducing the number of individual tokens alone in sequence labeling tasks will result in the loss of the syntactic relations of the sentence. This paper proposes a weighted variant of conditional random field (CRF) for data-imbalanced sequence labeling tasks. Such a weighting strategy can alleviate data distribution imbalances between majority and minority classes. Instead of using softmax in the output layer, the CRF can capture the relationship of labels between tokens. The locally interpretable model-agnostic explanations (LIME) algorithm was applied to investigate performance differences between models with and without the weighted loss function. Experimental results on two different ADR tasks show that the proposed model outperforms previously proposed sequence labeling methods. Post-marketing drug safety surveillance offers the chance to detect serious ADRs resulting in hospitalization and ADRs occurring in patients, e.g., patients with high comorbidity and receiving drugs that are administered only in hospitals. This monitoring has traditionally been accomplished by surveying users. Recently, the automatically recording ADR of users in social media can greatly help biopharmaceutical enterprises to improve their products. Previous methods of name entity recognition in natural language processing were usually performed on the corpora with a balanced data distribution. Conversely, the datasets for ADR detection are extremely imbalanced. As a result, the detector tends to ignore the ADR symptoms and the related indications, which are more important. In this study, we propose a weighted CRF model based on BERT for the detection task of ADR. A weighted variant of the Viterbi Algorithm is implemented to assign more weight to the minority class, forcing the model to pay more attention to minority classes to ensure effective detection. The results suggested that the proposed method provides a significant performance boost without changing the model architecture in imbalanced-data tasks.
DOI: 10.1109/tc.2018.2805683
发表时间: 2018-08-01
影响因子: 3.7
作者:
Kanduri, Anil;Haghbayan, Mohammad-Hashem;Liljeberg, Pasi
通讯作者: Liljeberg, Pasi
DOI: 10.1038/clpt.2013.47
发表时间: 2013-06
影响因子: 6.7
作者:
LePendu, P.;Iyer, S. V.;Bauer-Mehren, A.;Harpaz, R.;Mortensen, J. M.;Podchiyska, T.;Ferris, T. A.;Shah, N. H.
通讯作者: Shah, N. H.
DOI: 10.1613/jair.953
发表时间: 2002-01-01
影响因子: 5
作者:
Chawla, NV;Bowyer, KW;Kegelmeyer, WP
通讯作者: Kegelmeyer, WP
DOI: 10.1007/978-3-540-30115-8_7
发表时间: 2004-01-01
期刊: MACHINE LEARNING: ECML 2004, PROCEEDINGS
影响因子: --
作者:
Akbani, R;Kwek, S;Japkowicz, N
通讯作者: Japkowicz, N
DOI: 10.1109/tnn.2010.2066988
发表时间: 2010-10-01
影响因子: --
作者:
Chen, Sheng;He, Haibo;Garcia, Edwardo A.
通讯作者: Garcia, Edwardo A.