Extracting drug-drug interactions from literature using a rich feature-based linear kernel approach.

Extracting drug-drug interactions from literature using a rich feature-based linear kernel approach.
复制标题

DOI:
10.1016/j.jbi.2015.03.002
复制
发表时间:
2015-06
影响因子:
4.5
通讯作者:
Wilbur, W. John
Wilbur, W. John
中科院分区:
医学3区
文献类型:
--
作者:
Kim, Sun;Liu, Haibin;Yeganova, Lana;Wilbur, W. John

文献摘要

参考文献

被引文献

相似文献

识别未知的药物相互作用对于早期发现药物不良反应有很大好处。尽管存在多种药物相互作用 (DDI) 信息资源,但此类信息的财富被埋藏在呈指数级增长的非结构化医学文本中。这就需要开发文本挖掘技术来识别 DDI。最先进的 DDI 提取方法使用支持向量机 (SVM) 和非线性复合内核来探索文献中的不同上下文。虽然计算成本较低,但基于线性内核的系统在 DDI 提取任务中尚未实现可比的性能。在这项工作中,我们提出了一种使用线性内核来识别 DDI 信息的高效且可扩展的系统。所提出的方法包括两个步骤:识别 DDI 并将四种不同 DDI 类型之一分配给预测的药物对。我们证明,当配备一组丰富的词汇和句法特征时,线性 SVM 分类器能够在检测 DDI 方面实现具有竞争力的性能。此外,事实证明,一对一策略对于解决 DDI 类型分类中的不平衡问题至关重要。应用于 DDIExtraction 2013 语料库时,我们的系统获得了 0.670 的 F1 分数,而 DDIExtraction 2013 挑战赛中排名前两名的参赛团队的 F1 分数为 0.651 和 0.609,这两个团队都基于非线性核方法。
Identifying unknown drug interactions is of great benefit in the early detection of adverse drug reactions. Despite existence of several resources for drug-drug interaction (DDI) information, the wealth of such information is buried in a body of unstructured medical text which is growing exponentially. This calls for developing text mining techniques for identifying DDIs. The state-of-the-art DDI extraction methods use Support Vector Machines (SVMs) with non-linear composite kernels to explore diverse contexts in literature. While computationally less expensive, linear kernel-based systems have not achieved a comparable performance in DDI extraction tasks. In this work, we propose an efficient and scalable system using a linear kernel to identify DDI information. The proposed approach consists of two steps: identifying DDIs and assigning one of four different DDI types to the predicted drug pairs. We demonstrate that when equipped with a rich set of lexical and syntactic features, a linear SVM classifier is able to achieve a competitive performance in detecting DDIs. In addition, the one-against-one strategy proves vital for addressing an imbalance issue in DDI type classification. Applied to the DDIExtraction 2013 corpus, our system achieves an F1 score of 0.670, as compared to 0.651 and 0.609 reported by the top two participating teams in the DDIExtraction 2013 challenge, both based on non-linear kernel methods.
通过评估跨核心学习,全路径图核用于蛋白质 - 蛋白质相互作用提取。
DOI: 10.1186/1471-2105-9-s11-s2
发表时间: 2008-11-19
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Airola, Antti;Pyysalo, Sampo;Bjoerne, Jari;Pahikkala, Tapio;Ginter, Filip;Salakoski, Tapio
通讯作者: Salakoski, Tapio
使用单词和句法特征对蛋白质-蛋白质相互作用文章进行分类。
DOI: 10.1186/1471-2105-12-s8-s9
发表时间: 2011-10-03
期刊: BMC bioinformatics
影响因子: 3
作者:
Kim S;Wilbur WJ
通讯作者: Wilbur WJ
DOI: 10.1371/journal.pone.0060954
发表时间: 2013
期刊: PloS one
影响因子: 3.7
作者:
Liu H;Hunter L;Kešelj V;Verspoor K
通讯作者: Verspoor K
DOI: 10.1186/2041-1480-3-3
发表时间: 2012-04-01
影响因子: 1.9
作者:
Liu H;Christiansen T;Baumgartner WA Jr;Verspoor K
通讯作者: Verspoor K
DOI: 10.1093/bioinformatics/btl616
发表时间: 2007-02-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Fundel, Katrin;Kueffner, Robert;Zimmer, Ralf
通讯作者: Zimmer, Ralf