High throughput profile-profile based fold recognition for the entire human proteome

High throughput profile-profile based fold recognition for the entire human proteome
复制标题

DOI:
10.1186/1471-2105-7-288
复制
发表时间:
2006-06-07
期刊:
影响因子:
3
通讯作者:
Jones, David T.
Jones, David T.
中科院分区:
生物学4区
文献类型:
--
作者:
McGuffin, Liam J.;Smith, Richard T.;Jones, David T.

文献摘要

被引文献

相似文献

背景资料:为了维护最全面的结构注释数据库,我们必须使用最新的轮廓-轮廓折叠识别方法对每个蛋白质组进行定期更新。根据需要进行这些更新的能力是跟上序列和结构数据库的定期更新所必需的。提供最高质量的结构模型需要使用最新的可用序列数据库和折叠文库运行的最密集的轮廓-轮廓折叠识别方法。然而,运行这些方法在这样一个定期的基础上,每一个测序的蛋白质组需要大量的处理能力,在本文中,我们描述和基准的JYDE(作业产量分布环境)系统,这是一个元调度器,旨在工作在集群,如太阳网格引擎(SGE)或秃鹰。我们证明了JYDE的能力,分布在多个独立的网格域的基因组规模的折叠识别的负载。我们使用最新的配置文件,配置文件版本的mGenTHREADER软件,以注释最新版本的人类蛋白质组对最新的序列和结构databases.Results的最短的时间:我们表明,我们的JYDE系统能够扩展到大量的密集折叠识别作业运行在几个独立的计算机集群。使用我们的JYDE系统,我们已经能够注释99.9%的蛋白质序列内的人类蛋白质组在不到24小时内,利用超过500个CPU从3个独立的Grid domain.Conclusion:这项研究清楚地表明了按需进行高质量的结构注释的主要真核生物的蛋白质组的可行性。具体来说,我们已经表明,它现在可以提供完整的定期更新的配置文件,配置文件为基础的折叠识别模型,整个真核蛋白质组,通过使用网格中间件,如JYDE。
Background: In order to maintain the most comprehensive structural annotation databases we must carry out regular updates for each proteome using the latest profile-profile fold recognition methods. The ability to carry out these updates on demand is necessary to keep pace with the regular updates of sequence and structure databases. Providing the highest quality structural models requires the most intensive profile-profile fold recognition methods running with the very latest available sequence databases and fold libraries. However, running these methods on such a regular basis for every sequenced proteome requires large amounts of processing power.In this paper we describe and benchmark the JYDE (Job Yield Distribution Environment) system, which is a meta-scheduler designed to work above cluster schedulers, such as Sun Grid Engine (SGE) or Condor. We demonstrate the ability of JYDE to distribute the load of genomic-scale fold recognition across multiple independent Grid domains. We use the most recent profile-profile version of our mGenTHREADER software in order to annotate the latest version of the Human proteome against the latest sequence and structure databases in as short a time as possible.Results: We show that our JYDE system is able to scale to large numbers of intensive fold recognition jobs running across several independent computer clusters. Using our JYDE system we have been able to annotate 99.9% of the protein sequences within the Human proteome in less than 24 hours, by harnessing over 500 CPUs from 3 independent Grid domains.Conclusion: This study clearly demonstrates the feasibility of carrying out on demand high quality structural annotations for the proteomes of major eukaryotic organisms. Specifically, we have shown that it is now possible to provide complete regular updates of profile-profile based fold recognition models for entire eukaryotic proteomes, through the use of Grid middleware such as JYDE.