Extracting and characterizing gene-drug relationships from the literature

Extracting and characterizing gene-drug relationships from the literature
复制标题

DOI:
10.1097/00008571-200409000-00002
复制
发表时间:
2004-09-01
期刊:
PHARMACOGENETICS
影响因子:
--
通讯作者:
Altman, RB
Altman, RB
中科院分区:
其他
文献类型:
--
作者:
Chang, JT;Altman, RB

文献摘要

被引文献

相似文献

药物遗传学的一个基本任务是收集和分类基因和药物之间的关系。目前,这些有用的信息还没有在任何数据库中全面汇总,仍然分散在已出版的文献中。尽管有人工收集这些信息的努力,但他们受到已出版的关于基因-药物关系的文献规模的限制。因此,我们研究了从文献中提取和表征基因和药物之间的药物遗传关系的计算方法。我们首先评估了共现法在识别相关基因和药物方面的有效性。然后,我们使用有监督的机器学习算法,将来自药物遗传学和药物基因组学知识库(PharmGKB)的基因和药物之间的关系分类为五个类别,这些类别已被活跃的药物遗传学研究人员定义为与他们的工作相关。最终的共生算法能够从文献中提取78%的相关基因和药物,发表在一篇综述文章中。我们的算法随后将来自PharmGKB的基因和药物之间的关系分类为五类,准确率为74%。我们已经在http://bionlp.stanford.edu/genedrug/的一个补充网站上提供了这些数据,可以从文本中准确地提取基因-药物关系,并将其分类。尽管我们已经确定的关系没有捕捉到文献中经常做出的细节和细微的区别,但这些方法将帮助科学家追踪不断增长的文献,并创建信息资源来支持未来的发现。(C)2004年,里平科特·威廉姆斯·威尔金斯。
A fundamental task of pharmacogenetics is to collect and classify relationships between genes and drugs. Currently, this useful information has not been comprehensively aggregated in any database and remains scattered throughout the published literature. Although there are efforts to collect this information manually, they are limited by the size of the published literature on gene-drug relationships. Therefore, we investigated computational methods to extract and characterize pharmacogenetic relationships between genes and drugs from the literature. We first evaluated the effectiveness of the co-occurrence method in identifying related genes and drugs. We then used supervised machine learning algorithms to classify the relationships between genes and drugs from the Pharmacogenetics and Pharmacogenomics Knowledge Base (PharmGKB) into five categories that have been defined by active pharmacogenetic researchers as relevant to their work. The final co-occurrence algorithm was able to extract 78% of the related genes and drugs that were published in a review article from the literature. Our algorithm subsequently classified the relationships between genes and drugs from the PharmGKB into five categories with 74% accuracy. We have made the data available on a supplementary website at http://bionlp.stanford.edu/genedrug/ Gene-drug relationships can be accurately extracted from text and classified into categories. Although the relationships that we have identified do not capture the details and fine distinctions often made in the literature, these methods will help scientists to track the ever-growing literature and create information resources to support future discoveries. (C) 2004 Lippincott Williams Wilkins.