UTILIZING TERAGRID TO DETECT REMOTE SIMILARITY PROTEIN SEQUENCES
UTILIZING TERAGRID TO DETECT REMOTE SIMILARITY PROTEIN SEQUENCES
批准号:
7723381
负责人:
MARK FIENUP
金额:
$0.05万
依托单位国家:
美国
项目类别:
财政年份:
2008
资助国家:
美国
项目状态:
已结题
起止时间:
2008-08-01 至 2009-07-31
关键词:
AlgorithmsAmino Acid SequenceAmino Acid Sequence DatabasesAmino Acid Sequence HomologyAreaComputer Retrieval of Information on Scientific Projects DatabaseComputing MethodologiesDataDatabasesDepositionDetectionFundingGenomicsGoalsGrantHandHuman Genome ProjectInstitutionMethodsNuclear Magnetic ResonancePeptide Sequence DeterminationPhasePlayProcessProteinsResearchResearch PersonnelResourcesRoleRunningSourceStagingStructureTechniquesTestingTimeUnited States National Institutes of HealthX-Ray Crystallographycostprotein structureprotein structure function
中文摘要
点击翻译按钮获取中文摘要
英文摘要
This subproject is one of many research subprojects utilizing the
resources provided by a Center grant funded by NIH/NCRR. The subproject and
investigator (PI) may have received primary funding from another NIH source,
and thus could be represented in other CRISP entries. The institution listed is
for the Center, which is not necessarily the institution for the investigator.
The structure of a protein is often a key to its function. However, significant time and cost is required to determine the structure of a protein by experimental methods, such as the X-ray crystallography or the Nuclear Magnetic Resonance. There are currently less than 50,000 protein structures deposited in the Protein Data Bank (PDB), of which about 80% are redundant. On the other hand, the genomic sequencing efforts, such as the Human Genome Project, have populated protein sequence databases with well over 5 million sequences. With the increasing gap between known sequences and experimentally determined structures, the computational methods capable of predicting the structure and function of proteins will play an increasing role in protein annotation studies. The ultimate goal of the research described in this proposal is to develop a new protein sequence homology detection method that leverages the growing body of protein sequence data in ways that existing methods do not. The increased sensitivity in recognizing relationships between amino acid sequences will be achieved through the applications of intermediate sequence search strategies and profile-profile techniques. To date, the progress in this area has been limited by the lack of the computational resources needed to perform the transitive profile-profile search. We propose to utilize the TeraGrid to develop and test the first intermediate profile-profile algorithm for detecting protein sequence similarities. The algorithm constructs a sequential profile for the input amino acid sequence (target) and uses it to transitively search the database of all representative profiles for sequences in nr. In the transitive search, the matches found after running the first sequence comparison are used as new queries against the database. The whole process is repeated, iteratively with these new matches. The similarity between the target profile and the profile from the database is established through the intermediate sequences. Our project will be carried out in two stages: 1. In the first stage we will generate the set of representative alignment profiles for sequences from the non-redundant protein sequence database nr. 2. In the second phase we will deploy and test our algorithm.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
UTILIZING TERAGRID TO DETECT REMOTE SIMILARITY PROTEIN SEQUENCES
-
批准号:7956240
-
项目类别:
-
资助金额:$0.08万
-
财政年份:2009
-
负责人:MARK FIENUP
-
依托单位:
REMOTE PROTEIN SEQUENCE HOMOLOGY DETECTION
-
批准号:7956227
-
项目类别:
-
资助金额:$0.08万
-
财政年份:2009
-
负责人:MARK FIENUP
-
依托单位:
REMOTE PROTEIN SEQUENCE HOMOLOGY DETECTION
-
批准号:7723368
-
项目类别:
-
资助金额:$0.05万
-
财政年份:2008
-
负责人:MARK FIENUP
-
依托单位:
海外基金