Generic information can retrieve known biological associations: implications for biomedical knowledge discovery.

Generic information can retrieve known biological associations: implications for biomedical knowledge discovery.
复制标题

DOI:
10.1371/journal.pone.0078665
复制
发表时间:
2013
期刊:
影响因子:
3.7
通讯作者:
Schultes EA
Schultes EA
中科院分区:
综合性期刊3区
文献类型:
--
作者:
van Haagen HH;'t Hoen PA;Mons B;Schultes EA

文献摘要

参考文献

被引文献

相似文献

根据文本挖掘文献构建的加权语义网络可用于检索已知的蛋白质-蛋白质或基因-疾病关联,并且已被证明可以在文献中明确陈述之前几年预测关联。我们的文本挖掘系统可识别超过 640,000 个生物医学概念:一些是特定的(即基因或蛋白质的名称),另一些是通用的(例如“智人”)。通用概念可能在自动信息检索、提取和推理中发挥重要作用,但也可能导致概念过载,并将检索和推理与低相关性甚至虚假链接混淆。在这里,我们尝试通过从加权语义网络过滤通用概念(节点过滤)或通用概念链接(边缘过滤)来优化蛋白质-蛋白质相互作用(PPI)的检索性能。首先,我们根据网络属性定义了量化概念特异性的指标。然后使用这些指标,我们系统地从网络中过滤通用信息,同时监控已知蛋白质-蛋白质相互作用的检索性能。我们还系统地从网络中过滤特定信息(逆过滤),并评估仅由通用信息组成的网络的检索性能。过滤通用或特定信息会导致检索性能出现两阶段响应:最初过滤的影响很小,但超过临界阈值后网络性能会突然下降。与预期相反,仅由通用信息组成的网络表现出与也包含特定概念的未过滤网络相当的检索性能。此外,使用单个通用概念的分析表明它们可以有效支持已知蛋白质-蛋白质相互作用的检索。例如,概念“结合”表示 PPI 检索,概念“突变异常”表示基因-疾病关联。通用概念对于信息检索很重要,并且不能从语义网络中删除而不会对检索性能产生负面影响。
Weighted semantic networks built from text-mined literature can be used to retrieve known protein-protein or gene-disease associations, and have been shown to anticipate associations years before they are explicitly stated in the literature. Our text-mining system recognizes over 640,000 biomedical concepts: some are specific (i.e., names of genes or proteins) others generic (e.g., ‘Homo sapiens’). Generic concepts may play important roles in automated information retrieval, extraction, and inference but may also result in concept overload and confound retrieval and reasoning with low-relevance or even spurious links. Here, we attempted to optimize the retrieval performance for protein-protein interactions (PPI) by filtering generic concepts (node filtering) or links to generic concepts (edge filtering) from a weighted semantic network. First, we defined metrics based on network properties that quantify the specificity of concepts. Then using these metrics, we systematically filtered generic information from the network while monitoring retrieval performance of known protein-protein interactions. We also systematically filtered specific information from the network (inverse filtering), and assessed the retrieval performance of networks composed of generic information alone. Filtering generic or specific information induced a two-phase response in retrieval performance: initially the effects of filtering were minimal but beyond a critical threshold network performance suddenly drops. Contrary to expectations, networks composed exclusively of generic information demonstrated retrieval performance comparable to unfiltered networks that also contain specific concepts. Furthermore, an analysis using individual generic concepts demonstrated that they can effectively support the retrieval of known protein-protein interactions. For instance the concept “binding” is indicative for PPI retrieval and the concept “mutation abnormality” is indicative for gene-disease associations. Generic concepts are important for information retrieval and cannot be removed from semantic networks without negative impact on retrieval performance.
DOI: 10.1093/nar/gkm881
发表时间: 2008-01
影响因子: 14.9
作者:
Bruford, Elspeth A.;Lush, Michael J.;Wright, Mathew W.;Sneddon, Tam P.;Povey, Sue;Birney, Ewan
通讯作者: Birney, Ewan
DOI: 10.1016/j.ijmedinf.2007.07.004
发表时间: 2008-05-01
影响因子: 4.9
作者:
Jelier, Rob;Schuemie, Martijn J.;Kors, Jan A.
通讯作者: Kors, Jan A.
DOI: 10.1093/nar/gkq1237
发表时间: 2011-01
影响因子: 14.9
作者:
Maglott D;Ostell J;Pruitt KD;Tatusova T
通讯作者: Tatusova T
DOI: 10.1140/epjst/e2012-01692-1
发表时间: 2012-11-01
影响因子: 2.8
作者:
van Harmelen, F.;Kampis, G.;Helbing, D.
通讯作者: Helbing, D.
DOI: 10.1038/nrg2918
发表时间: 2011-01
期刊: Nature reviews. Genetics
影响因子: --
作者:
通讯作者: --