A probabilistic model for mining implicit 'chemical compound-gene' relations from literature

A probabilistic model for mining implicit 'chemical compound-gene' relations from literature
复制标题

DOI:
10.1093/bioinformatics/bti1141
复制
发表时间:
2005-01
期刊:
影响因子:
5.8
通讯作者:
Shanfeng Zhu;Y. Okuno;G. Tsujimoto;Hiroshi Mamitsuka
Shanfeng Zhu;Y. Okuno;G. Tsujimoto;Hiroshi Mamitsuka
中科院分区:
生物学3区
文献类型:
--
作者:
Shanfeng Zhu;Y. Okuno;G. Tsujimoto;Hiroshi Mamitsuka

文献摘要

被引文献

相似文献

动机 化合物的重要性在分子生物学中得到了更多的强调,“化学基因组学”近年来引起了广泛的关注。因此,当前分子生物学的一个重要问题是识别与生物相关的化合物(更具体地说,药物)和基因。文献中生物实体的共现是一种简单、全面且流行的寻找这些实体关联的技术。我们的重点是从文献中的共现中挖掘隐含的“化合物和基因”关系。结果我们提出了一种概率模型,称为混合方面模型(MAM),以及一种估计其参数的算法,以有效地同时处理不同类型的共现数据集。我们不仅通过使用 MEDLINE 记录生成的数据进行交叉验证,还通过使用 ChEBI 数据库中化学化合物和基因之间关系的独立的人工数据集进行测试来检查我们方法的性能。在这两种情况下,我们对三种不同类型的共现数据集(即复合-基因、基因-基因和复合-复合共现)进行了实验。实验结果表明,由所有数据集训练的 MAM 优于由其他数据集组合训练的任何简单模型,并且在所有情况下差异均具有统计显着性。特别是,我们发现合并复合复合共现对于提高预测性能最有效。我们最终使用我们的方法计算了所有未知化合物-基因(更具体地说,药物-基因)对的可能性,并根据可能性选择了前 20 对。我们从生物学、医学和制药的角度验证了它们。
MOTIVATION The importance of chemical compounds has been emphasized more in molecular biology, and 'chemical genomics' has attracted a great deal of attention in recent years. Thus an important issue in current molecular biology is to identify biological-related chemical compounds (more specifically, drugs) and genes. Co-occurrence of biological entities in the literature is a simple, comprehensive and popular technique to find the association of these entities. Our focus is to mine implicit 'chemical compound and gene' relations from the co-occurrence in the literature. RESULTS We propose a probabilistic model, called the mixture aspect model (MAM), and an algorithm for estimating its parameters to efficiently handle different types of co-occurrence datasets at once. We examined the performance of our approach not only by a cross-validation using the data generated from the MEDLINE records but also by a test using an independent human-curated dataset of the relationships between chemical compounds and genes in the ChEBI database. We performed experimentation on three different types of co-occurrence datasets (i.e. compound-gene, gene-gene and compound-compound co-occurrences) in both cases. Experimental results have shown that MAM trained by all datasets outperformed any simple model trained by other combinations of datasets with the difference being statistically significant in all cases. In particular, we found that incorporating compound-compound co-occurrences is the most effective in improving the predictive performance. We finally computed the likelihoods of all unknown compound-gene (more specifically, drug-gene) pairs using our approach and selected the top 20 pairs according to the likelihoods. We validated them from biological, medical and pharmaceutical viewpoints.