UTILIZING TERAGRID TO DETECT REMOTE SIMILARITY PROTEIN SEQUENCES
UTILIZING TERAGRID TO DETECT REMOTE SIMILARITY PROTEIN SEQUENCES
批准号:
7956240
负责人:
MARK FIENUP
金额:
$0.08万
依托单位国家:
美国
项目类别:
财政年份:
2009
资助国家:
美国
项目状态:
已结题
起止时间:
2009-08-01 至 2010-07-31
关键词:
AlgorithmsAmino Acid SequenceAmino Acid Sequence DatabasesAmino Acid Sequence HomologyAreaBiomedical ResearchComputer Retrieval of Information on Scientific Projects DatabaseComputing MethodologiesDataDatabasesDepositionDetectionFundingGenomicsGoalsGrantHandHigh Performance ComputingHuman Genome ProjectInstitutionMethodsNuclear Magnetic ResonancePeptide Sequence DeterminationPhasePlayProcessProteinsResearchResearch PersonnelResourcesRoleRunningSourceStagingStructureTechniquesTestingTimeUnited States National Institutes of HealthX-Ray Crystallographycomputing resourcescostprotein structureprotein structure function
中文摘要
这个子项目是许多研究子项目中利用
资源由NIH/NCRR资助的中心拨款提供。子项目和
调查员(PI)可能从NIH的另一个来源获得了主要资金,
并因此可以在其他清晰的条目中表示。列出的机构是
该中心不一定是调查人员的机构。
蛋白质的结构往往是其功能的关键。然而,通过X射线结晶学或核磁共振等实验方法来确定蛋白质的结构需要大量的时间和成本。目前蛋白质数据库(PDB)中保存的蛋白质结构不到5万个,其中约80%是冗余的。另一方面,基因组测序工作,如人类基因组计划,已经用500多万个序列填充了蛋白质序列数据库。随着已知序列和实验确定的结构之间的差距越来越大,能够预测蛋白质结构和功能的计算方法将在蛋白质注释研究中发挥越来越大的作用。本提案中描述的研究的最终目标是开发一种新的蛋白质序列同源性检测方法,该方法能够以现有方法所不具备的方式利用不断增长的蛋白质序列数据。通过应用中间序列搜索策略和剖面分析技术,将提高识别氨基酸序列之间关系的灵敏度。到目前为止,这一领域的进展一直受到缺乏执行传递性简档-简档搜索所需的计算资源的限制。我们建议利用TeraGrid来开发和测试第一个用于检测蛋白质序列相似性的中间轮廓-轮廓算法。该算法为输入的氨基酸序列(目标)构建一个序列图谱,并使用它在所有代表性图谱的数据库中传递地搜索nr中的序列。在传递性搜索中,运行第一次序列比较后找到的匹配项被用作对数据库的新查询。对于这些新的匹配,整个过程重复进行。通过中间序列建立目标简档和来自数据库的简档之间的相似性。我们的项目将分两个阶段进行:1.在第一阶段,我们将从非冗余蛋白质序列数据库nr中为序列生成一组具有代表性的比对轮廓。2.在第二阶段,我们将部署和测试我们的算法。
英文摘要
This subproject is one of many research subprojects utilizing the
resources provided by a Center grant funded by NIH/NCRR. The subproject and
investigator (PI) may have received primary funding from another NIH source,
and thus could be represented in other CRISP entries. The institution listed is
for the Center, which is not necessarily the institution for the investigator.
The structure of a protein is often a key to its function. However, significant time and cost is required to determine the structure of a protein by experimental methods, such as the X-ray crystallography or the Nuclear Magnetic Resonance. There are currently less than 50,000 protein structures deposited in the Protein Data Bank (PDB), of which about 80% are redundant. On the other hand, the genomic sequencing efforts, such as the Human Genome Project, have populated protein sequence databases with well over 5 million sequences. With the increasing gap between known sequences and experimentally determined structures, the computational methods capable of predicting the structure and function of proteins will play an increasing role in protein annotation studies. The ultimate goal of the research described in this proposal is to develop a new protein sequence homology detection method that leverages the growing body of protein sequence data in ways that existing methods do not. The increased sensitivity in recognizing relationships between amino acid sequences will be achieved through the applications of intermediate sequence search strategies and profile-profile techniques. To date, the progress in this area has been limited by the lack of the computational resources needed to perform the transitive profile-profile search. We propose to utilize the TeraGrid to develop and test the first intermediate profile-profile algorithm for detecting protein sequence similarities. The algorithm constructs a sequential profile for the input amino acid sequence (target) and uses it to transitively search the database of all representative profiles for sequences in nr. In the transitive search, the matches found after running the first sequence comparison are used as new queries against the database. The whole process is repeated, iteratively with these new matches. The similarity between the target profile and the profile from the database is established through the intermediate sequences. Our project will be carried out in two stages: 1. In the first stage we will generate the set of representative alignment profiles for sequences from the non-redundant protein sequence database nr. 2. In the second phase we will deploy and test our algorithm.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
REMOTE PROTEIN SEQUENCE HOMOLOGY DETECTION
-
批准号:7956227
-
项目类别:
-
资助金额:$0.08万
-
财政年份:2009
-
负责人:MARK FIENUP
-
依托单位:
UTILIZING TERAGRID TO DETECT REMOTE SIMILARITY PROTEIN SEQUENCES
-
批准号:7723381
-
项目类别:
-
资助金额:$0.05万
-
财政年份:2008
-
负责人:MARK FIENUP
-
依托单位:
REMOTE PROTEIN SEQUENCE HOMOLOGY DETECTION
-
批准号:7723368
-
项目类别:
-
资助金额:$0.05万
-
财政年份:2008
-
负责人:MARK FIENUP
-
依托单位:
海外基金