Advanced correlation analyses to infer sequence and structural determinants of protein function
Advanced correlation analyses to infer sequence and structural determinants of protein function
批准号:
10093067
负责人:
ANDREW F NEUWALD
金额:
$30.9万
依托单位国家:
美国
项目类别:
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-02-01 至 2023-01-31
关键词:
AcetyltransferaseBase SequenceBayesian AnalysisBenchmarkingBiochemicalBiologicalBiomedical ResearchCaringCatalysisCodeCollaborationsCommon CoreCorrelation StudiesCouplingDNA Repair EndonucleaseDataData SetDatabasesDependenceDimerizationDrug DesignEnsureEvaluationFeedbackFormulationGoalsGuanosine Triphosphate PhosphohydrolasesHealthHumanHydrogen BondingIndividualInvestigationJointsLinkMeasuresMediatingMethodsMindModelingMolecularMolecular BiologyPatternPerformancePhosphoric Monoester HydrolasesPositioning AttributeProcessPropertyProtein EngineeringProteinsQuality ControlReliability of ResultsResearchResearch PersonnelRoleSamplingSensitivity and SpecificitySequence AlignmentSpecificitySpeedStatistical ModelsStructureSubgroupSystemTestingTimeValidationbasefollow-uphuman diseaseimprovedinnovationinositol-1,4,5-trisphosphate 5-phosphataseinsightinterestmemberopen sourcepersonalized medicineprogramsprotein functionthree dimensional structuretool
中文摘要
点击翻译按钮获取中文摘要
英文摘要
PROJECT SUMMARY
A long-term goal of molecular biology is assigning functional and mechanistic roles to specific protein residues,
beyond the obvious roles in catalysis. Although this task is hindered by the relative sparsity of experimentally-
based sequence annotations, it is facilitated by an abundance of sequence data augmented by structural data.
This has spurred sequence- and structure-based prediction of function determining residues using a wide
variety of methods. However, by focusing on experimentally characterized functions, these methods disfavor
recognition of residues involved in important uncharacterized functions, insofar as these will be benchmarked
incorrectly as false positives. Instead, this project focuses more generally on inferring functionally-relevant
residues (FRRs) by allowing the sequence data itself to reveal its most statistically surprising properties
without making assumptions about what will be found. We argue that, in the absence of experimental
annotations, it is only possible to directly link individual residues to other residues and such residue sets to
structural features. This project will make such associations by identifying sequence-to-sequence and
sequence-to-structure correlations, and will focus solely on the observed data rather than on predicting
(unseen) biochemical properties. The goal is to obtain hypothesis-generating observations for experimental
follow up. Aim 1 will create advanced tools for characterizing correlated residue patterns due to functional
divergence with each pattern consisting of an arbitrary number of residues. Aim 2 will develop a tool to
probabilistically assess correlations between independent sequence- and structurally-defined residue sets.
This tool will be modified for other purposes, including the evaluation of FRR-prediction programs. Aim 3 will
integrate Aims 1 & 2 methods and direct coupling analysis (DCA) into a nearly comprehensive system for
sequence/structural correlation analysis. (Unlike the correlations under Aims 1 & 2, DCA focuses on direct
correlations between residue pairs.) This strategy involves a high degree of model complexity and optimization
over diverse sequence properties synergistically (due to interrelationships and dependencies) and over
alternative models and parameters; hence, considerable care is required to ensure reliable results. Therefore,
we will apply information theoretical principles to adjust accurately for multiple hypotheses, to avoid under- and
over-fitting to the data, and to eliminate inherent biases. Aim 3 will also characterize the relationships among
the various types of correlations. We will apply these tools to large, functionally diverse superfamilies in
collaboration with researchers interested in these proteins. Using tools developed under Aim 2 and hundreds
of conserved domain datasets, Aim 4 will rigorously benchmark the performance of tools developed under
Aims 1 & 3 relative to competing methods. This project will aid research efforts in protein engineering, the
molecular basis of human disease, drug design and personalized medicine.
期刊论文(15)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1186/s12859-021-04108-5
发表时间:
2021-04-15
期刊:
BMC bioinformatics
影响因子:
3
作者:
[Chen X, Shi X, Neuwald AF, Hilakivi-Clarke L, Clarke R, Xuan J]
通讯作者:
Xuan J
DOI:
10.1038/s41598-020-79603-5
发表时间:
2021-01-11
期刊:
Scientific reports
影响因子:
4.6
作者:
[Chen X, Gu J, Neuwald AF, Hilakivi-Clarke L, Clarke R, Xuan J]
通讯作者:
Xuan J
DOI:
10.3390/ijms23063024
发表时间:
2022-03-11
期刊:
International journal of molecular sciences
影响因子:
5.6
作者:
[Chapoval SP, Lee M, Lemmer A, Ajayi O, Qi X, Neuwald AF, Keegan AD]
通讯作者:
Keegan AD
IntAPT: integrated assembly of phenotype-specific transcripts from multiple RNA-seq profiles.
IntAPT:来自多个 RNA-seq 配置文件的表型特异性转录本的集成组装。
DOI:
10.1093/bioinformatics/btaa852
发表时间:
2021
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
作者:
[Shi,Xu, Neuwald,AndrewF, Wang,Xiao, Wang,Tian-Li, Hilakivi-Clarke,Leena, Clarke,Robert, Xuan,Jianhua]
通讯作者:
Xuan,Jianhua
DOI:
10.1016/j.csbj.2022.04.005
发表时间:
2022
期刊:
Computational and structural biotechnology journal
影响因子:
6
作者:
[]
通讯作者:
共 10 条
Predicting common protein mechanisms by the light of evolution
-
批准号:7258353
-
项目类别:
-
资助金额:$0.0万
-
财政年份:2006
-
负责人:ANDREW F NEUWALD
-
依托单位:
Predicting common protein mechanisms by the light of evolution
-
批准号:7651998
-
项目类别:
-
资助金额:$24.59万
-
财政年份:2006
-
负责人:ANDREW F NEUWALD
-
依托单位:
Predicting common protein mechanisms by the light of evolution
-
批准号:7471672
-
项目类别:
-
资助金额:$19.52万
-
财政年份:2006
-
负责人:ANDREW F NEUWALD
-
依托单位:
Predicting common protein mechanisms by the light of evolution
-
批准号:7683169
-
项目类别:
-
资助金额:$25.49万
-
财政年份:2006
-
负责人:ANDREW F NEUWALD
-
依托单位:
Predicting common protein mechanisms by the light of evolution
-
批准号:7138450
-
项目类别:
-
资助金额:$8.39万
-
财政年份:2006
-
负责人:ANDREW F NEUWALD
-
依托单位:
Predicting common protein mechanisms by the light of evolution
-
批准号:7497470
-
项目类别:
-
资助金额:$25.49万
-
财政年份:2006
-
负责人:ANDREW F NEUWALD
-
依托单位:
Advanced Sequence-Based Prediction of Protein Function
-
批准号:6559851
-
项目类别:
-
资助金额:$15.2万
-
财政年份:1998
-
负责人:ANDREW F NEUWALD
-
依托单位:
Advanced Sequence-Based Prediction of Protein Function
-
批准号:6584899
-
项目类别:
-
资助金额:$4.76万
-
财政年份:1998
-
负责人:ANDREW F NEUWALD
-
依托单位:
Advanced Sequence-Based Prediction of Protein Function
-
批准号:6796747
-
项目类别:
-
资助金额:$37.35万
-
财政年份:1998
-
负责人:ANDREW F NEUWALD
-
依托单位:
Advanced Sequence-Based Prediction of Protein Function
-
批准号:6650870
-
项目类别:
-
资助金额:$37.35万
-
财政年份:1998
-
负责人:ANDREW F NEUWALD
-
依托单位:
Advanced Sequence-Based Prediction of Protein Function
-
批准号:6528388
-
项目类别:
-
资助金额:$37.35万
-
财政年份:1998
-
负责人:ANDREW F NEUWALD
-
依托单位:
Advanced Sequence-Based Prediction of Protein Function
-
批准号:6399662
-
项目类别:
-
资助金额:$22.19万
-
财政年份:1998
-
负责人:ANDREW F NEUWALD
-
依托单位:
Advanced Sequence-Based Prediction of Protein Function
-
批准号:7394278
-
项目类别:
-
资助金额:$3.49万
-
财政年份:1998
-
负责人:ANDREW F NEUWALD
-
依托单位:
Advanced Sequence-Based Prediction of Protein Function
-
批准号:6941253
-
项目类别:
-
资助金额:$33.86万
-
财政年份:1998
-
负责人:ANDREW F NEUWALD
-
依托单位:
ADVANCED SEQUENCE BASED PREDICTION OF PROTEIN FUNCTION
-
批准号:2897404
-
项目类别:
-
资助金额:$33.49万
-
财政年份:1998
-
负责人:ANDREW F NEUWALD
-
依托单位:
ADVANCED SEQUENCE BASED PREDICTION OF PROTEIN FUNCTION
-
批准号:6185230
-
项目类别:
-
资助金额:$34.48万
-
财政年份:1998
-
负责人:ANDREW F NEUWALD
-
依托单位:
ADVANCED SEQUENCE BASED PREDICTION OF PROTEIN FUNCTION
-
批准号:2740163
-
项目类别:
-
资助金额:$36.06万
-
财政年份:1998
-
负责人:ANDREW F NEUWALD
-
依托单位:
海外基金