COMPUTATIONAL LEARNING & DISCOVERY FOR BIOLOGICAL SEQUENCE, STRUCTURE, FUNCTION
COMPUTATIONAL LEARNING & DISCOVERY FOR BIOLOGICAL SEQUENCE, STRUCTURE, FUNCTION
批准号:
7723019
负责人:
RAJ REDDY
金额:
$0.07万
依托单位国家:
美国
项目类别:
财政年份:
2008
资助国家:
美国
项目状态:
已结题
起止时间:
2008-06-01 至 2009-05-31
关键词:
AlgorithmsAmino Acid SequenceAreaBenchmarkingBiologicalBiological Neural NetworksClassificationComputer Retrieval of Information on Scientific Projects DatabaseData SetEpidermal Growth Factor ReceptorFundingG-Protein-Coupled ReceptorsGene ProteinsGenomeGoalsGrantInstitutionLanguageLearningLinguisticsMetabotropic Glutamate ReceptorsModelingPathway interactionsPeptide Sequence DeterminationPeptidesPerformancePostdoctoral FellowProcessProtein FamilyProtocols documentationResearchResearch PersonnelResourcesRhodopsinSchemeSourceStructureStudentsSystemTestingTrainingUnited States National Institutes of HealthViralVirus DiseasesWorkbacteriophage tailspike proteinbasedesignprotein foldingprotein protein interactionprotein structureresearch study
中文摘要
点击翻译按钮获取中文摘要
英文摘要
This subproject is one of many research subprojects utilizing the
resources provided by a Center grant funded by NIH/NCRR. The subproject and
investigator (PI) may have received primary funding from another NIH source,
and thus could be represented in other CRISP entries. The institution listed is
for the Center, which is not necessarily the institution for the investigator.
Seven focus areas in the realm of protein structure have been identified for application of the language analogy approach. These focus areas are: protein folding, conformational changes, protein-protein interactions, protein/gene networks and pathways, secondary structure and repetitive folds prediction and segmentation, protein family classification, and genome comparison. The ultimate goal is to develop linguistic models for each that are capable of advancing the understanding of these areas. The protocol followed in this process consists of several steps. The first step is to utilize existing "benchmark" datasets or to define datasets suitable for training and testing of these models. As controls, existing approaches in the focus areas, if available, are studied and a scheme is designed for evaluating the language model approaches and comparing them to existing other approaches. The next step is to implement our language approach. This implementation initially needs to meet one or both of two requirements: (i) the system has to perform equally well or better than existing systems as defined in step 2 and/or (ii) it needs to provide interpretable biological hypotheses. For example, a neural network might be the algorithm with best performance in a classification task, but the underlying features resulting in this performance can be unclear. A language-based approach that might have lesser performance but allows the researcher to analyze the types of features that result in successful classification can be used to build hypotheses on the fundamental building blocks of protein sequence language. The final step in the protocol is to design and carry out experiments that specifically test these hypotheses. The following systems have been chosen as experimental test cases for the language models: G protein coupled receptors (GPCR) such as rhodopsin, metabotropic glutamate receptors, epidermal growth factor receptor, viral tailspike protein, virus infection process, peptide n-grams. For each of the seven focus areas, we are working to identify or develop benchmark datasets for training and testing of linguistic models. Students and postdoctoral fellows participate in all aspects of the projects.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
COMPUTATIONAL LEARNING & DISCOVERY FOR BIOLOGICAL SEQUENCE, STRUCTURE, FUNCTION
-
批准号:7602013
-
项目类别:
-
资助金额:$0.11万
-
财政年份:2007
-
负责人:RAJ REDDY
-
依托单位:
COMPUTATIONAL LEARNING & DISCOVERY FOR BIOLOGICAL SEQUENCE, STRUCTURE, FUNCTION
-
批准号:7369285
-
项目类别:
-
资助金额:$0.12万
-
财政年份:2006
-
负责人:RAJ REDDY
-
依托单位:
COMPUTATIONAL LEARNING & DISCOVERY FOR BIOLOGICAL SEQUENCE, STRUCTURE, FUNCTION
-
批准号:7182240
-
项目类别:
-
资助金额:$0.12万
-
财政年份:2005
-
负责人:RAJ REDDY
-
依托单位:
COMPUTATIONAL LEARNING & DISCOVERY FOR BIOLOGICAL SEQUENCE, STRUCTURE, FUNCTION
-
批准号:6978546
-
项目类别:
-
资助金额:$0.24万
-
财政年份:2004
-
负责人:RAJ REDDY
-
依托单位:
海外基金