课题基金 / 基金详情

项目摘要

项目成果

CHITTIBABU GUDA的其他基金

相似基金

相关文献

中文摘要
翻译
描述(由申请人提供):蛋白质亚细胞定位的精确知识在系统生物学研究中非常重要,因为大多数细胞过程在细胞中受到空间限制。这种空间背景对于更好地理解与跨越亚细胞边界的疾病途径相关的细胞内串扰和细胞信号传导中所涉及的蛋白质的各种作用至关重要。实验确定的定位仅可用于UniProt数据库中约1%的蛋白质。计算方法可以补充实验工作,以确定许多未知定位的蛋白质的定位。现有的计算方法具有有限的范围和适用性,因此不适合于蛋白质组范围内的定位预测。此外,由于缺乏任何实验验证,这些预测的可靠性值得怀疑。在这个项目中,我们建议开发一个全面的系统,使我们能够创建准确和全面的目录的亚细胞和亚细胞器蛋白质组的所有测序基因组的动物物种。这个系统是基于我们最近发表的计算方法,称为ngmix,使用“n-gram”肽(蛋白质的固定长度序列)来建立准确的贝叶斯模型,用于亚细胞和亚细胞器类别的分类。此外,nginx非常适合蛋白质组范围的预测和预测定位于多个细胞器的蛋白质。基于nggram方法,我们提出了一种新的方法,通过使用先进的计算概念,如半监督学习,分层贝叶斯分类和集成方法,并通过实施替换矩阵比较n-gram同源性。所有这些方法都在其他领域取得了成功,因此预计将大大提高我们方法的准确性。一组400人的蛋白质,其定位由我们的新方法预测将在正常和癌细胞系的人进行实验测试,使用GFP融合和表达,然后在共聚焦显微镜下可视化。这一步将使我们能够确定我们的方法在每个细胞器的每个分数阈值的预测准确性。使用最佳评分阈值,将进行蛋白质组范围的预测,并为所有测序的动物物种基因组生成实验已知和预测的亚细胞和亚细胞器蛋白质组的详细目录。此外,将开发一个用于改进方法的独立软件包,并根据通用公共许可证(GPL)向研究界发布。将开发一个在线网络服务器,以便进行在线预测,并使人们能够查阅编目数据和本项目制作的软件。总之,拟议的综合系统将提供一个“黄金标准”的实验建立的本地化数据集,一种新的预测方法,预测本地化的实验验证,以及一个公共网络服务器来预测或访问数据集和软件工具在这个项目中开发。这些资源将被证明是非常有价值的生物医学研究界在推进系统生物学研究的许多方面。
英文摘要
DESCRIPTION (provided by applicant): Precise knowledge of the subcellular localization of proteins is very important in systems biology research because most cellular processes are spatially constrained in the cell. This spatial context is essential to gain a better understanding of the various roles of proteins involved in the intra-cellular cross-talk and cell signaling associated with disease pathways that span across subcellular boundaries. Experimentally-determined localizations are available only for about 1% of the proteins in the UniProt database. Computational methods can complement experimental efforts in determining the localization of many proteins with unknown localization. Existing computational methods have limited scope and applicability, and hence are not suitable for proteome-wide prediction of localizations. Moreover, the reliability of these predictions is questionable due to lack of any experimental validation. In this project, we propose the development of a comprehensive system that will enable us to create accurate and comprehensive catalogs of subcellular and suborganellar proteomes of all sequenced genomes of animal species. This system is based on our recently published computational method known as ngLOC, that uses 'n-gram' peptides (fixed-length subsequences of proteins) to build accurate Bayesian models for classification of subcellular and suborganellar classes. Additionally, ngLOC is well suited for proteome-wide predictions and to predict proteins localized to multiple organelles. Based on the ngLOC approach, we propose to develop a new method by using advanced computational concepts such as semi-supervised learning, hierarchical Bayesian classification and ensemble approaches, and by implementing substitutions matrices to compare n-gram homology. All of these methods have proven success in other domains and hence are expected to substantially improve the accuracy of our method. A set of 400 human proteins whose localizations are predicted by our new method will be experimentally tested in normal and cancer cell lines of human, using GFP-fusion and expression followed by visualization under confocal microscope. This step would allow us to determine the prediction accuracy of our method at each score threshold for each organelle. Using optimal score thresholds, proteome-wide predictions will be carried out and detailed catalogs of experimentally-known and predicted subcellular and suborganellar proteomes will be generated for all sequenced genomes of animal species. Additionally, a standalone software package for the improved method will be developed and released to the research community under the General Public License (GPL). An online web server will be developed to make predictions online, and to enable access to the cataloged data and to the software produced in this project. In summary, the proposed comprehensive system will deliver a 'gold-standard' dataset of experimentally established localizations, a novel methodology for prediction, experimental validation of predicted localizations, and a public web server to predict or to access datasets and the software tool developed in this project. These resources will prove to be very valuable to the biomedical research community in advancing the many facets of systems biology research.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Biomedical Informatics, Bioinformatics, and Cyberinfrastructure Enhancement Core
Biomedical Informatics, Bioinformatics, and Cyberinfrastructure Enhancement Core
Cataloging the subcellular and suborganellar proteomes of sequenced genomes
Cataloging the subcellular and suborganellar proteomes of sequenced genomes
海外基金