Algorithms for effective querying of compound graph-based pathway databases

Algorithms for effective querying of compound graph-based pathway databases
复制标题

DOI:
10.1186/1471-2105-10-376
复制
发表时间:
2009-11-16
期刊:
影响因子:
3
通讯作者:
Babur, Ozgun
Babur, Ozgun
中科院分区:
生物学4区
文献类型:
--
作者:
Dogrusoz, Ugur;Cetintas, Ahmet;Babur, Ozgun

文献摘要

被引文献

相似文献

背景:基于图的通路本体和数据库被广泛用于表示关于细胞过程的数据。这种表示使得有可能以编程方式集成细胞网络,并使用图论的概念来研究它们,以预测它们的结构和动态特性。这种图表示的扩展,即分层结构或复合图,其中生物网络的成员可以递归地包含某种逻辑上相似的生物对象组的子网络,为生物途径的分析提供了许多额外的好处,包括通过分解成不同的组件或模块来降低复杂性。在这方面,有必要有效地查询这种集成的大型复合网络,以便在高效算法和软件工具的帮助下提取感兴趣的子网络。为了实现这一目标,我们开发了一个查询框架,沿着一些图论算法,从简单的邻域查询到最短路径到反馈循环,适用于各种基于图的路径数据库,从PPI(蛋白质-蛋白质相互作用)到代谢和信号通路。该框架是独特的,因为它可以解释复合或嵌套结构和普遍存在的实体中存在的途径数据。此外,查询可以通过“AND”和“OR”运算符彼此相关,并且可以递归地组织成树,其中一个查询的结果可以是另一个查询的源和/或目标,以形成更复杂的查询。该算法的查询组件内的一个新版本的软件工具PATIKAweb(路径分析工具的集成和知识获取),并已被证明是有用的回答了一些生物学上的重大问题,大型基于图形的pathway databases.Conclusion:PATIKA项目的网站是http://www.patika.org PATIKAweb版本2.1可在http://web.patika.org。
Background: Graph-based pathway ontologies and databases are widely used to represent data about cellular processes. This representation makes it possible to programmatically integrate cellular networks and to investigate them using the well-understood concepts of graph theory in order to predict their structural and dynamic properties. An extension of this graph representation, namely hierarchically structured or compound graphs, in which a member of a biological network may recursively contain a sub-network of a somehow logically similar group of biological objects, provides many additional benefits for analysis of biological pathways, including reduction of complexity by decomposition into distinct components or modules. In this regard, it is essential to effectively query such integrated large compound networks to extract the sub-networks of interest with the help of efficient algorithms and software tools.Results: Towards this goal, we developed a querying framework, along with a number of graph-theoretic algorithms from simple neighborhood queries to shortest paths to feedback loops, that is applicable to all sorts of graph-based pathway databases, from PPIs (protein-protein interactions) to metabolic and signaling pathways. The framework is unique in that it can account for compound or nested structures and ubiquitous entities present in the pathway data. In addition, the queries may be related to each other through "AND" and "OR" operators, and can be recursively organized into a tree, in which the result of one query might be a source and/or target for another, to form more complex queries. The algorithms were implemented within the querying component of a new version of the software tool PATIKAweb (Pathway Analysis Tool for Integration and Knowledge Acquisition) and have proven useful for answering a number of biologically significant questions for large graph-based pathway databases.Conclusion: The PATIKA Project Web site is http://www.patika.org. PATIKAweb version 2.1 is available at http://web.patika.org.