课题基金 / 基金详情

A STUDY OF THE INTEROPERABILITY BETWEEN TERAGRID AND CNGRID BY EXPERIMENTING RE

A STUDY OF THE INTEROPERABILITY BETWEEN TERAGRID AND CNGRID BY EXPERIMENTING RE
通过RE实验研究TERAGRID和CNGRID之间的互操作性
批准号:
7601478
负责人:
JINBO XU
金额:
$0.03万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2007
资助国家:
美国
项目状态:
已结题
起止时间:
2007-08-01 至 2008-07-31

项目摘要

项目成果

JINBO XU的其他基金

相似基金

相关文献

中文摘要
翻译
这个子项目是许多研究子项目中利用 资源由NIH/NCRR资助的中心拨款提供。子项目和 调查员(PI)可能从NIH的另一个来源获得了主要资金, 并因此可以在其他清晰的条目中表示。列出的机构是 该中心不一定是调查人员的机构。 通过实验真实规模的生物信息学应用程序研究TeraGrid和CNGrid之间的互操作性1.简介TeraGrid被认为是支持美国开放科学研究的世界上最大的网格基础设施。目前部署的TeraGrid为多学科的科学和教育社区提供了极大的计算和存储容量[1]。中国国家网格作为中国第一阶段(2002年-2005年)网格技术研发的国家级试验平台,目标是在第二阶段(2007年-2010年)成为支持开放式e-Science活动的生产环境。它是由中国政府根据高科技863计划赞助的,一期1,300万美元,二期3,600万美元。目前的CNGrid环境是围绕8个20Tflops/S、200TB存储的全国超算中心互联互通而构建的。这两个网格试验台的互操作有利于双边协作科学应用,包括更大规模的共享资源池、不同底层软件堆栈的统一访问接口以及潜在的更好的服务质量。2.目的本项目的主要目的是通过实验运行在TeraGrid和CNGrid上的真实大小的生物信息学应用程序来研究这两个平台之间的互操作性问题和解决方案。从这个项目中获得的经验将作为CNGrid互操作性活动的基础,这是下一阶段CNGrid软件的子任务之一[3]。此外,该项目还可以为未来全球主流网格项目的互操作做出贡献,这在OGF被称为GIN[4]。交付成果将停留在两个层面:a)在应用层面,我们将探索在典型网格应用的国际测试平台上可以获得哪些好处;b)在中间件层面,我们将为常见服务互联互通的技术问题提供概念验证解决方案,包括身份验证和授权、数据管理、作业提交、资源发现和监控等。3.任务两个生物信息学应用程序Raptor[5]和Treeback[6]将首次作为评估两个试验床互操作性的基准应用程序。Raptor是最好的蛋白质结构预测程序之一。TreePack(以前称为SCATD)是一个基于蛋白质骨架结构树分解的侧链预测程序。它们都可以在SMP机器、集群系统和大规模并行处理平台上运行。猛禽将通过一个门户网站公开供生物科学家使用,该门户网站为网格试验台提供前端服务。在从初步实验中获得足够的经验后,该项目的PI和Co-PI将与CNGrid应用伙伴合作,开展大规模的生物信息学研究活动。TeraGrid使用CTSS作为中间件,它由Globus工具包(v2和v4)、秃鹰和其他实用程序组成。CNGrid有自己的软件堆栈,名为GOS[3],它采用了一种面向服务的方法,符合许多开放标准,包括WS-I基本概要、WS-Security和SAML。实际上,CNGrid软件将重点放在VO级别的管理服务上,这些服务可以将TeraGrid站点作为其VOS的成员。具体地说,该项目可以按照预期的时间表分为以下任务:(1)在选定的TeraGrid和CNGrid站点上部署应用程序,并调查应用程序如何在这两个软件堆栈(CTSS和CNGrid GOS)上执行。(M1)(2)开发应用程序级网关(例如,通过CNGrid门户或桌面应用程序),允许CNGrid节点基于CNGrid软件客户端库但使用TeraGrid帐户访问TeraGrid生物信息学资源。(M2)(3)CNGrid的CA加入PMA等国际PKI联合会。探索一种与TeraGrid站点上的CNGrid证书相关的身份验证和授权方法。成功将生物信息学作业提交并监控到具有CNGrid证书的TeraGrid站点。(M3)(4)进行步骤2、3的相反部分,以允许从TeraGrid访问CNGrid资源(M4)(5)对从两个试验床跨站点运行的应用的性能和开销进行一系列基准测试,识别和整合国际网格试验床的要求、好处和问题。(M5-M6)4.资源请求的合理性执行并行生物信息学实验应用所需的5,000个SU。10G存储用于应用程序以及基因组数据库,10G暂存空间用于测试数据。5.参考文献[1]TeraGrid,http://www.teragrid.org/[2]CNGrid,http://www.cngrid.org/[3]X Xie,N肖,Z Xu,L查,W Li,H Yu,CNGrid Software 2:面向服务的网格计算方法,英国e-Science全体会议论文集,2005-allhands.org.uk[4]gin,http://forge.gridforum.org/sf/go/projects.gin/wiki[5]徐锦波,徐颖,金东苏,李明.《猛禽:基于线性规划的最优蛋白质线索分析》,创刊,生物信息学与计算生物学,2003年4月[6]徐劲波。基于树状分解的蛋白质侧链快速包装。RECOMB 2005。
英文摘要
This subproject is one of many research subprojects utilizing the resources provided by a Center grant funded by NIH/NCRR. The subproject and investigator (PI) may have received primary funding from another NIH source, and thus could be represented in other CRISP entries. The institution listed is for the Center, which is not necessarily the institution for the investigator. A Study of the Interoperability between TeraGrid and CNGrid by Experimenting Real-size Bioinformatics Applications 1. Introduction TeraGrid is known as the world's largest grid infrastructure for supporting open scientific research in the US. Current deployment of TeraGrid provides extremely large computing and storage capacity for multi-disciplinary scientific and education communities [1]. The China National Grid (CNGrid) [2], serving as a nation-scale testbed for grid technology research and development in China in its first phase(2002-2005), aims to become a production environment for supporting open e-Science activities in its second phase(2007-2010). It is sponsored by Chinese government under the Hi-tech 863 program, with an award of $13 million for first-phase and $36 million for second-phase. Current CNGrid environment has been built around the interconnection of eight national-wide supercomputing centers with a capacity of 20Tflops/s and 200TB storage. Interoperation of these two grid testbeds is beneficial to bi-lateral collaborative scientific applications, in terms of larger scale of pooled resources for sharing, uniform access interfaces for disparate underlying software stacks, and potential better quality of service. 2. Objective The main purpose of this project is to study interoperability issues and solutions between TeraGrid and CNGrid by experimenting real-size bioinformatics applications that run cross these two testbeds. Experiences gained from this project serve as a basis for CNGrid Interoperability Activity, one of sub tasks of next-phase CNGrid software [3]. Furthermore, this project could also contribute to future interoperating of world-wide main-stream grid projects, which is coined as GIN [4] at OGF. The deliverable results will reside on two levels: a) At the application level, we will explore what benefits could be gained on international testbeds for typical grid applications; b) At the middleware level, we will give proof-of-concept solution of technical issues regarding interconnection of common services, including authentication and authorization, data management, job submission, resource discovery and monitoring, and so on. 3. Tasks Two bioinformatics applications, RAPTOR [5] and Treeback [6], will be firstly taken as benchmarking applications for evaluating the interoperability of two testbeds. RAPTOR is one of the best protein structure prediction program. TreePack (called SCATD before) is a side-chain prediction program based on tree-decomposition of a protein backbone structure. Both of them can run on a SMP machine, a cluster system and a massively parallel processing platform. The RAPTOR will be made public available for use by bio-scientists via a web portal that serving a front-end to grid testbeds. After gaining enough experiences from initial experiments, PI and Co-PI of this project will work together with CNGrid application partners, to conduct large-scale bioinformatics research activities. TeraGrid employs CTSS as the middleware which consists of Globus Toolkit (v2 and v4), Condor and other utilities. CNGrid has its own software stack, named GOS [3], which adopts a service oriented approach compliant to many open standards including WS-I basic profile, WS-Security and SAML. Actually CNGrid software puts focus on VO-level management services that could take TeraGrid sites as members of its VOs. Concretely, this project could be divided into the following tasks with an expected schedule: (1) Application deployment on both selected TeraGrid and CNGrid sites and investigating how applications are executed on both software stacks ( CTSS and CNGrid GOS). (M1) (2) Develop application-level gateway (e.g. via a CNGrid portal or a desktop application) that allows CNGrid nodes to access TeraGrid bioinformatics resources, based on CNGrid software client libraries but using TeraGrid account. ( M2 ) (3) CAs of CNGrid join international PKI federations like PMA. Explore an approach for authentication and authorization with respect to CNGrid certificates on TeraGrid sites. Successfully submit and monitor bioinformatics jobs to TeraGrid sites with CNGrid certificates. (M3) (4) Doing the vice versa part of step 2,3 , to allow the access of CNGrid resources from TeraGrid (M4) (5) Conducting a series of benchmark tests on performance and overhead of applications run cross sites both from two testbeds, identify and consolidate requirements, benefits and issues for international grid testbeds. (M5-M6) 4. Justification for Resource Requests 5,000 SUs needed for executing parallel bioinformatics experimental applications. 10G storage for applications together with genome databases, 10G scratch space for test data. 5. References [1] TeraGrid, http://www.teragrid.org/ [2] CNGrid, http://www.cngrid.org/ [3] X Xie, N Xiao, Z Xu, L Zha, W Li, H Yu, CNGrid Software 2: Service Oriented Approach to Grid Computing, the proceedings of the UK e-Science All Hands Meeting, 2005 - allhands.org.uk [4] GIN, http://forge.gridforum.org/sf/go/projects.gin/wiki [5] Jinbo Xu, Ying Xu, Dongsup Kim, Ming Li. RAPTOR: Optimal Protein Threading by Linear Programming, the inaugural issue, Journal of Bioinformatics and Computational Biology, April 2003 [6] Jinbo Xu. Rapid Protein Side-Chain Packing via Tree Decomposition. RECOMB 2005.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
New Computational Methods for Data-driven Protein Structure Prediction
New Computational Methods for Data-driven Protein Structure Prediction
New Computational Methods for Data-driven Protein Structure Prediction
New Computational Methods for Data-driven Protein Structure Prediction
海外基金