课题基金 / 基金详情

项目摘要

项目成果

Steven E Brenner的其他基金

相似基金

相关文献

中文摘要
翻译
描述(申请人提供):基因组和超基因组计划揭示了数百万蛋白质的遗传序列,其生物学解释需要了解它们的功能。预测蛋白质功能最成功的方法之一是将所有可用的功能数据进化关系整合到一个协调的系统发育树中。这种被称为系统基因组学的方法被誉为高度准确和概念上的优雅,但它的应用受到了对领域专家艰苦分析的精致依赖的限制。我们将加强、评估和应用一种统计方法,利用系统基因组学原理预测蛋白质的功能。我们的方法被称为SIFTER(通过进化关系对函数的统计推断),目前作为一个原型存在。在这个建议中,我们将增强核心算法,以考虑到领域架构,使其方法变得更一致的统计,并适应更大范围的蛋白质可能的功能。我们将改进分子进化模型的关键内部参数,并提高结果的可解释性。我们将使程序能够接受更多典型的蛋白质序列进行分析,并使用更广泛的信息(包括数据库注释、序列和结构基序)作为功能的证据。最终,Siefter将能够将其他功能预测方法整合到其系统发育背景中。筛子的性能将使用经过充分研究的家庭进行严格评估。我们将与主要的蛋白质数据库合作,部署Screter,用于蛋白质标注的中等规模应用。实验验证对于真正测试Screter的性能至关重要,巧合的是,它丰富了我们对几个蛋白质家族的生物学理解。我们将使用SIFTER对Nudex蛋白进行最佳选择,以进行实验表征。除了分析这些蛋白质外,我们还将通过结构基因组学中心对蛋白质的分子功能进行盲目预测,然后我们将对提供给我们的有前途的候选蛋白质进行生化表征。完整的筛选系统应该会对目前的蛋白质功能预测方法提供重大改进,这几乎与所有分子生物学家都有直接关系。通过解锁基因组序列中编码的蛋白质功能信息,这项工作对公共卫生的意义是明确和直接的。这些方法将使人们能够了解人类和模型生物中涉及疾病和健康所必需的蛋白质。筛子的应用还将允许详细了解病原体和共生微生物区系的蛋白质。这些方法将为进一步研究通过基因组计划确定的任何蛋白质奠定基础。
英文摘要
DESCRIPTION (provided by applicant): Genome and metagenome projects have revealed the genetic sequence of millions of proteins, whose biological interpretation requires understanding of their function. One of the most successful approaches for predicting proteins' functions is the integration of all available functional data evolutionary relationships in a reconciled phylogenetic tree. This method, known as phylogenomics, has been heralded as highly accurate and conceptually elegant, but its application has been limited by its exquisite dependency upon painstaking analyses by domain experts. We will enhance, assess, and apply a statistical method for predicting protein function using phylogenomic principles. Our approach, known as SIFTER (Statistical Inference of Function Through Evolutionary Relationships) presently exists as a prototype. In this proposal, we will enhance the core algorithms to take account of domain architecture, to become more consistently statistical in its approach, and to accommodate a larger range of possible functions for proteins. We will improve the key internal parameters of the molecular evolution model, and improve interpretability of the results. We will make the program capable of accepting more typical protein sequences for analysis, and of using a wider range of information (including database annotations, sequence & structure motifs) as evidence of function. Ultimately, SIFTER will be capable of incorporating other function prediction approaches within its phylogenetic context. The performance of SIFTER will be rigorously assessed using well-studied families. We will collaborate with major protein databases to deploy SIFTER for medium-scale application in protein annotation. Experimental validation will be essential to truly test SIFTER'S performance and, coincidentally, enrich our biological understanding of several protein families. We will use SIFTER to make an optimal selection of Nudix proteins for experimental characterization. In addition to assaying these proteins, we will also make blind predictions of molecular function of proteins being characterized by structural genomics centers, and we will then biochemically characterize promising candidate proteins provided to us. The completed SIFTER system should provide a significant improvement over current approaches for protein function prediction, of direct relevance to nearly all molecular biologists. The significance of this work for public health is clear and immediate, by unlocking protein function information encoded in genome sequences. These methods will allow understanding of proteins implicated in disease and necessary for health, in humans as well as model organisms. Application of SIFTER will also permit detailed understanding of pathogens' and commensal microbiota's proteins. These methods will be a foundation for the further study of any protein identified through genome projects.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Identification of Candidate Disease-Causing Variants
Informatics Infrastructure and Bioinformatics Analysis
Identification of Candidate Disease-Causing Variants
Identification of Candidate Disease-Causing Variants
海外基金