springD2A: capturing uncertainty in disease-drug association prediction with model integration

springD2A: capturing uncertainty in disease-drug association prediction with model integration
复制标题

DOI:
10.1093/bioinformatics/btab820
复制
发表时间:
2021-12
期刊:
影响因子:
5.8
通讯作者:
Weiwen Wang;Xiwen Zhang;D. Dai
Weiwen Wang;Xiwen Zhang;D. Dai
中科院分区:
生物学3区
文献类型:
--
作者:
Weiwen Wang;Xiwen Zhang;D. Dai

文献摘要

被引文献

相似文献

动机 旨在为现有药物寻找新适应症的药物重定位一直是药物研发的一种有效策略。在我们仅将已确认的疾病 - 药物关联作为正例对的情况下,以往研究通常从未知的疾病 - 药物对中构建一组疾病 - 药物负例对(我们不知道药物和疾病是否相关),以训练用于疾病 - 药物关联预测(药物重定位)的模型。这些负例对中的药物和疾病可能存在潜在关联,但大多数研究都忽略了它们。 结果 我们提出了一种方法springD2A,用于捕捉负例对中的不确定性,并区分正例对和未知对,因为前者更可靠。在springD2A中,我们对负例对的损失引入了一种类似弹簧的惩罚机制,如果它们在单位球面上过于接近,惩罚力度就大,如果距离适中,惩罚则较轻。我们还设计了一种顺序采样方法,其中一个未知的疾病 - 药物对被采样为负例的概率与其被预测为正例的得分成正比。在顺序采样过程中学习多个模型,并且我们采用基于参数和基于特征的集成方案来提高性能。实验表明springD2A是一种用于药物重定位的有效工具。 可用性 springD2A的Python实现以及本研究中使用的数据集可在https://github.com/wangyuanhao/springD2A获取。 补充信息 补充数据可在Bioinformatics在线获取。
MOTIVATION Drug repositioning that aims to find new indications for existing drugs has been an efficient strategy for drug discovery. In the scenario where we only have confirmed disease-drug associations as positive pairs, a negative set of disease-drug pairs is usually constructed from the unknown disease-drug pairs in previous studies, where we do not know whether drugs and diseases can be associated, to train a model for disease-drug association prediction (drug repositioning). Drugs and diseases in these negative pairs can potentially be associated, but most studies have ignored them. RESULTS We present a method, springD2A, to capture the uncertainty in the negative pairs, and to discriminate between positive and unknown pairs because the former are more reliable. In springD2A, we introduce a spring-like penalty for the loss of negative pairs, which is strong if they are too close in a unit sphere, but mild if they are at a moderate distance. We also design a sequential sampling in which the probability of an unknown disease-drug pair sampled as negative is proportional to its score predicted as positive. Multiple models are learned during sequential sampling, and we adopt parameter- and feature-based ensemble schemes to boost performance. Experiments show springspringD2A is an effective tool for drug-repositioning. AVAILABILITY A python implementation of springD2A and datasets used in this study are available at https://github.com/wangyuanhao/springD2A. SUPPLEMENTARY INFORMATION Supplementary data are available at Bioinformatics online.