New developments on the cheminformatics open workflow environment CDK-Taverna

New developments on the cheminformatics open workflow environment CDK-Taverna
复制标题

DOI:
10.1186/1758-2946-3-54
复制
发表时间:
2011-12-13
影响因子:
8.6
通讯作者:
Steinbeck, Christoph
Steinbeck, Christoph
中科院分区:
化学2区
文献类型:
--
作者:
Truszkowski, Andreas;Jayaseelan, Kalai Vanii;Steinbeck, Christoph

文献摘要

被引文献

相似文献

背景资料:小分子的计算处理和分析是化学信息学和结构生物信息学的核心,也是它们在生物信息学中的应用。G.代谢组学或药物发现。流水线或工作流工具允许将I/O模块和算法的类似Lego(TM)的图形组装成一个复杂的工作流,该工作流可以很容易地部署、修改和测试,而无需将其实现为单片应用程序。CDK-Taverna项目旨在通过结合Taverna、化学开发工具包(CDK)或怀卡托知识分析环境(WEKA)等不同的开源项目,构建一个免费的开源化学信息学管道解决方案。CDK-Taverna的第一个集成版本1.0最近向公众发布。结果:CDK-Taverna项目迁移到其基础软件库的最新版本,并对其工人的架构进行了完整的重新设计(版本2.0)。64-现在支持位计算和多核线程的使用,以允许快速的存储器内处理和分析大量分子。以前的缺陷,如迭代数据阅读的解决方法被删除。组合化学相关的反应计数功能大大增强。实现用于计算小分子的天然产物相似性分数的附加功能以识别可能的候选药物。最后,数据分析功能通过新的工人进行扩展,这些工人提供对开源WEKA库的访问,用于集群和机器学习以及训练和测试集划分。新功能概述与usage scenaries.Conclusions:CDK-Taverna 2.0作为一个开源的化学信息学工作流解决方案成熟,成为一个免费提供的,越来越强大的工具,为生物科学。新的CDK-Taverna工作人员系列与活跃的Taverna社区开发并发布在myexperiment.org上的现有工作流程相结合,使分子科学家能够快速计算,处理和分析分子数据,例如在今天的系统生物学场景中通常会发现。
Background: The computational processing and analysis of small molecules is at heart of cheminformatics and structural bioinformatics and their application in e. g. metabolomics or drug discovery. Pipelining or workflow tools allow for the Lego (TM)-like, graphical assembly of I/O modules and algorithms into a complex workflow which can be easily deployed, modified and tested without the hassle of implementing it into a monolithic application. The CDK-Taverna project aims at building a free open-source cheminformatics pipelining solution through combination of different open-source projects such as Taverna, the Chemistry Development Kit (CDK) or the Waikato Environment for Knowledge Analysis (WEKA). A first integrated version 1.0 of CDK-Taverna was recently released to the public.Results: The CDK-Taverna project was migrated to the most up-to-date versions of its foundational software libraries with a complete re-engineering of its worker's architecture (version 2.0). 64-bit computing and multi-core usage by paralleled threads are now supported to allow for fast in-memory processing and analysis of large sets of molecules. Earlier deficiencies like workarounds for iterative data reading are removed. The combinatorial chemistry related reaction enumeration features are considerably enhanced. Additional functionality for calculating a natural product likeness score for small molecules is implemented to identify possible drug candidates. Finally the data analysis capabilities are extended with new workers that provide access to the open-source WEKA library for clustering and machine learning as well as training and test set partitioning. The new features are outlined with usage scenarios.Conclusions: CDK-Taverna 2.0 as an open-source cheminformatics workflow solution matured to become a freely available and increasingly powerful tool for the biosciences. The combination of the new CDK-Taverna worker family with the already available workflows developed by a lively Taverna community and published on myexperiment.org enables molecular scientists to quickly calculate, process and analyse molecular data as typically found in e.g. today's systems biology scenarios.