Statistical methods for higher order dependences to understand protein functions
Statistical methods for higher order dependences to understand protein functions
批准号:
10492723
负责人:
Wen Zhou
金额:
$20.33万
依托单位国家:
美国
项目类别:
财政年份:
2021
资助国家:
美国
项目状态:
已结题
起止时间:
2021-09-23 至 2024-08-31
关键词:
AlgorithmsAmino Acid SequenceAmino AcidsBiologicalCellsCharacteristicsComplexDataData SetDatabasesDependenceDrug DesignMeasuresMethodsModelingMolecularMolecular StructureOutcomePhysical environmentPrognosisProtein DynamicsProteinsPublic HealthScienceStatistical MethodsStructureSystemUncertaintyamino groupbasecomputer frameworkdiscrete datagenomic dataimprovednovelnovel strategiesprotein functionprotein structurestatisticsweb site
中文摘要
这项建议汇集了来自分子科学和统计学的强大团队,以解决重要的
如何在复杂系统中整合蛋白质结构和序列信息的问题。其中一些
这些数据最重要的特征是隐藏在其中的强相关性,与
序列数据中的成对相关性已经被常规用于预测结构接触。这里,
我们正在开发新的方法来使用海量数据集来提取更高阶的依赖关系,现在
从基因组学获得大量的序列数据是可能的;此外,在
在蛋白质结构中可以直接观察到这样的高阶相关性的分子结构,其中
氨基酸基团直接相互作用。重要的是,这些高阶依赖关系反映了密集的
单元格中需要适当统计特征的物理环境。一款免费的新车型
引入信息论方法对高阶相依关系进行量化,作为高阶相依关系的度量。
这个项目的中心方法。通过确定基于以下方面的统计推断的主要挑战
这一措施,我们开发、评估和改进了一个新的统计推断和计算框架
用于分析高阶相关性,具有一般类型的离散数据,受蛋白质的激励
多个序列数据。新的计算效率高的框架使发现可靠的
高阶依赖于量化不确定性的能力。这里的初步数据结合了
来自序列和结构的信息以产生意外结果,这些结果直接与
蛋白质结构的动力学。其结果是一种全新的方法来处理大容量
蛋白质序列数据和其他组学数据现已可用,大量数据即将到达
组学分析师的家门口。
英文摘要
This proposal brings together a strong team from molecular science and statistics to tackle the important
problem of how to integrate protein structure and sequence information in complex systems. Some of the
most important characteristics of these data are the strong correlations buried within them, with the
pairwise correlations in the sequence data already being routinely used to predict structural contacts. Here,
we are developing novel ways to use huge data sets to extract higher-order dependences, which are now
possible with the availability of the large volumes of sequence data from genomics; and in addition, in the
molecular structures such higher-order dependences are directly observable in the protein structures where
groups of amino acids interact directly. Importantly, these higher-order dependences reflect the dense
physical environment in the cell that requires for proper statistical characterization. A new model free
information-theoretic measure is introduced to quantify the higher-order dependences, which serves as the
central method in this project. By identifying the major challenges in drawing statistical inference based on
this measure, we develop, evaluate, and improve a new statistical inference and computational framework
for analyses of higher-order dependences with discrete data of a general type, motivated by the protein
multiple sequence data. The new computationally efficient framework makes it possible to discover reliable
higher-order dependences with the ability of quantifying uncertainty. The preliminary data here combine the
information from sequences and structures to yield unexpected results that immediately relate to the
dynamics of the protein structures. The outcome is an entirely new approach to handle the large volumes
of protein sequence data and other omics data now available and the enormous volumes about to arrive on
the doorsteps of omics analysts.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
ConProject-001
-
批准号:10707337
-
项目类别:
-
资助金额:$8.61万
-
财政年份:2021
-
负责人:Wen Zhou
-
依托单位:
ConProject-001
-
批准号:10492724
-
项目类别:
-
资助金额:$10.35万
-
财政年份:2021
-
负责人:Wen Zhou
-
依托单位:
Statistical methods for higher order dependences to understand protein functions
-
批准号:10378307
-
项目类别:
-
资助金额:$23.1万
-
财政年份:2021
-
负责人:Wen Zhou
-
依托单位:
Statistical methods for higher order dependences to understand protein functions
-
批准号:10707332
-
项目类别:
-
资助金额:$16.57万
-
财政年份:2021
-
负责人:Wen Zhou
-
依托单位:
ConProject-002
-
批准号:10492725
-
项目类别:
-
资助金额:$9.99万
-
财政年份:2021
-
负责人:Wen Zhou
-
依托单位:
ConProject-002
-
批准号:10707338
-
项目类别:
-
资助金额:$7.96万
-
财政年份:2021
-
负责人:Wen Zhou
-
依托单位:
海外基金