UTILIZING TERAGRID TO DETECT REMOTE SIMILARITY PROTEIN SEQUENCES
UTILIZING TERAGRID TO DETECT REMOTE SIMILARITY PROTEIN SEQUENCES
批准号:
7956240
负责人:
MARK FIENUP
金额:
$0.08万
依托单位国家:
美国
项目类别:
财政年份:
2009
资助国家:
美国
项目状态:
已结题
起止时间:
2009-08-01 至 2010-07-31
关键词:
AlgorithmsAmino Acid SequenceAmino Acid Sequence DatabasesAmino Acid Sequence HomologyAreaBiomedical ResearchComputer Retrieval of Information on Scientific Projects DatabaseComputing MethodologiesDataDatabasesDepositionDetectionFundingGenomicsGoalsGrantHandHigh Performance ComputingHuman Genome ProjectInstitutionMethodsNuclear Magnetic ResonancePeptide Sequence DeterminationPhasePlayProcessProteinsResearchResearch PersonnelResourcesRoleRunningSourceStagingStructureTechniquesTestingTimeUnited States National Institutes of HealthX-Ray Crystallographycomputing resourcescostprotein structureprotein structure function
中文摘要
该子项目是利用
由NIH/NCRR资助的中心赠款提供的资源。子项目和
研究者(PI)可能从另一个NIH来源获得主要资金,
因此可以在其他CRISP条目中表示。所列机构为
研究中心,而研究中心不一定是研究者所在的机构。
蛋白质的结构往往是其功能的关键。然而,通过实验方法,如X射线晶体学或核磁共振,确定蛋白质的结构需要大量的时间和成本。蛋白质数据库(Protein Data Bank,PDB)中保存了不到50,000个蛋白质结构,其中约80%是冗余的。另一方面,基因组测序的努力,如人类基因组计划,已填充蛋白质序列数据库远远超过500万个序列。随着已知序列与实验确定结构之间的差距越来越大,能够预测蛋白质结构和功能的计算方法将在蛋白质注释研究中发挥越来越大的作用。本提案中描述的研究的最终目标是开发一种新的蛋白质序列同源性检测方法,该方法以现有方法所不具备的方式利用不断增长的蛋白质序列数据。通过应用中间序列搜索策略和轮廓-轮廓技术,将实现识别氨基酸序列之间关系的增加的灵敏度。到目前为止,在这方面的进展一直受到限制,缺乏所需的计算资源进行传递的配置文件配置文件搜索。我们建议利用TeraGrid开发和测试的第一个中间配置文件配置文件算法检测蛋白质序列的相似性。该算法为输入的氨基酸序列(目标)构建了一个连续的轮廓,并使用它来传递搜索所有代表性轮廓的数据库中的nr中的序列。在传递性搜索中,在运行第一个序列比较之后找到的匹配被用作针对数据库的新查询。整个过程被重复,迭代地与这些新的匹配。通过中间序列建立目标谱与来自数据库的谱之间的相似性。我们的项目将分两个阶段进行:1.在第一阶段中,我们将为来自非冗余蛋白质序列数据库nr的序列生成一组代表性比对谱。2.在第二阶段,我们将部署和测试我们的算法。
英文摘要
This subproject is one of many research subprojects utilizing the
resources provided by a Center grant funded by NIH/NCRR. The subproject and
investigator (PI) may have received primary funding from another NIH source,
and thus could be represented in other CRISP entries. The institution listed is
for the Center, which is not necessarily the institution for the investigator.
The structure of a protein is often a key to its function. However, significant time and cost is required to determine the structure of a protein by experimental methods, such as the X-ray crystallography or the Nuclear Magnetic Resonance. There are currently less than 50,000 protein structures deposited in the Protein Data Bank (PDB), of which about 80% are redundant. On the other hand, the genomic sequencing efforts, such as the Human Genome Project, have populated protein sequence databases with well over 5 million sequences. With the increasing gap between known sequences and experimentally determined structures, the computational methods capable of predicting the structure and function of proteins will play an increasing role in protein annotation studies. The ultimate goal of the research described in this proposal is to develop a new protein sequence homology detection method that leverages the growing body of protein sequence data in ways that existing methods do not. The increased sensitivity in recognizing relationships between amino acid sequences will be achieved through the applications of intermediate sequence search strategies and profile-profile techniques. To date, the progress in this area has been limited by the lack of the computational resources needed to perform the transitive profile-profile search. We propose to utilize the TeraGrid to develop and test the first intermediate profile-profile algorithm for detecting protein sequence similarities. The algorithm constructs a sequential profile for the input amino acid sequence (target) and uses it to transitively search the database of all representative profiles for sequences in nr. In the transitive search, the matches found after running the first sequence comparison are used as new queries against the database. The whole process is repeated, iteratively with these new matches. The similarity between the target profile and the profile from the database is established through the intermediate sequences. Our project will be carried out in two stages: 1. In the first stage we will generate the set of representative alignment profiles for sequences from the non-redundant protein sequence database nr. 2. In the second phase we will deploy and test our algorithm.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
REMOTE PROTEIN SEQUENCE HOMOLOGY DETECTION
-
批准号:7956227
-
项目类别:
-
资助金额:$0.08万
-
财政年份:2009
-
负责人:MARK FIENUP
-
依托单位:
UTILIZING TERAGRID TO DETECT REMOTE SIMILARITY PROTEIN SEQUENCES
-
批准号:7723381
-
项目类别:
-
资助金额:$0.05万
-
财政年份:2008
-
负责人:MARK FIENUP
-
依托单位:
REMOTE PROTEIN SEQUENCE HOMOLOGY DETECTION
-
批准号:7723368
-
项目类别:
-
资助金额:$0.05万
-
财政年份:2008
-
负责人:MARK FIENUP
-
依托单位:
海外基金