Cataloging the subcellular and suborganellar proteomes of sequenced genomes
Cataloging the subcellular and suborganellar proteomes of sequenced genomes
批准号:
7918788
负责人:
CHITTIBABU GUDA
金额:
$0.0万
依托单位国家:
美国
项目类别:
财政年份:
2009
资助国家:
美国
项目状态:
已结题
起止时间:
2009-09-01 至 2010-09-02
关键词:
Amino Acid SequenceAnimalsAntibodiesAttentionBiologyBiomedical ResearchCancer cell lineCatalogingCatalogsCell physiologyCellsClassificationCommunitiesComplementComputer softwareComputing MethodologiesConfocal MicroscopyCytoplasmDataData SetDatabasesDevelopmentDisease PathwayEnsureFluorescenceGeneral PopulationGenomeGoldHumanHuman Cell LineImageryInfectionInternetKnowledgeLabelLearningLengthLicensingLocationMethodologyMethodsMicroscopeModelingOntologyOrganellesOutcomePeptide Sequence DeterminationPeptidesPlayProteinsProteomeProteomicsPublishingResearchResearch PersonnelResourcesRestRoleSignal TransductionSoftware ToolsSolutionsSourceSystemSystems BiologyTestingValidationbaseexpectationexpression vectorgenome sequencinghuman diseaseimprovednovelopen sourceprogramssoftware developmentsuccess
中文摘要
描述(由申请人提供):精确了解蛋白质的亚细胞定位在系统生物学研究中非常重要,因为大多数细胞过程在细胞中受到空间限制。这种空间背景对于更好地理解与跨越亚细胞边界的疾病途径相关的细胞内串扰和细胞信号传导所涉及的蛋白质的各种作用至关重要。实验确定的定位仅适用于UniProt数据库中约1%的蛋白质。计算方法可以补充实验工作,确定许多未知定位的蛋白质的定位。现有的计算方法范围和适用性有限,因此不适合蛋白质组范围的定位预测。此外,由于缺乏任何实验验证,这些预测的可靠性值得怀疑。在这个项目中,我们建议开发一个全面的系统,使我们能够创建准确和全面的亚细胞和亚细胞器蛋白质组目录的所有测序的动物物种的基因组。该系统基于我们最近发表的计算方法ngLOC,该方法使用“n-gram”肽(固定长度的蛋白质子序列)来构建精确的贝叶斯模型,用于亚细胞和亚细胞器分类。此外,ngLOC非常适合于蛋白质组范围的预测和预测定位于多个细胞器的蛋白质。基于ngLOC方法,我们提出了一种新的方法,通过使用先进的计算概念,如半监督学习,层次贝叶斯分类和集成方法,并通过实现替换矩阵来比较n-gram同源性。所有这些方法已经在其他领域被证明是成功的,因此有望大大提高我们的方法的准确性。我们的新方法预测了400种人类蛋白的定位,将在人类正常和癌细胞系中进行实验测试,使用gfp融合和表达,然后在共聚焦显微镜下显示。这一步将使我们能够确定我们的方法在每个细胞器的每个分数阈值下的预测精度。使用最佳评分阈值,将进行蛋白质组范围的预测,并为所有动物物种的测序基因组生成实验已知和预测的亚细胞和亚细胞器蛋白质组的详细目录。此外,改进方法的独立软件包将根据通用公共许可证(GPL)开发并发布给研究社区。将开发一个联机网络服务器,以便联机进行预测,并使人们能够访问已编目的数据和本项目编制的软件。总之,拟议的综合系统将提供实验建立的本地化的“金标准”数据集,一种新的预测方法,预测本地化的实验验证,以及一个公共web服务器来预测或访问数据集和本项目开发的软件工具。这些资源将被证明是非常宝贵的生物医学研究界在推进系统生物学研究的许多方面。
英文摘要
DESCRIPTION (provided by applicant): Precise knowledge of the subcellular localization of proteins is very important in systems biology research because most cellular processes are spatially constrained in the cell. This spatial context is essential to gain a better understanding of the various roles of proteins involved in the intra-cellular cross-talk and cell signaling associated with disease pathways that span across subcellular boundaries. Experimentally-determined localizations are available only for about 1% of the proteins in the UniProt database. Computational methods can complement experimental efforts in determining the localization of many proteins with unknown localization. Existing computational methods have limited scope and applicability, and hence are not suitable for proteome-wide prediction of localizations. Moreover, the reliability of these predictions is questionable due to lack of any experimental validation. In this project, we propose the development of a comprehensive system that will enable us to create accurate and comprehensive catalogs of subcellular and suborganellar proteomes of all sequenced genomes of animal species. This system is based on our recently published computational method known as ngLOC, that uses 'n-gram' peptides (fixed-length subsequences of proteins) to build accurate Bayesian models for classification of subcellular and suborganellar classes. Additionally, ngLOC is well suited for proteome-wide predictions and to predict proteins localized to multiple organelles. Based on the ngLOC approach, we propose to develop a new method by using advanced computational concepts such as semi-supervised learning, hierarchical Bayesian classification and ensemble approaches, and by implementing substitutions matrices to compare n-gram homology. All of these methods have proven success in other domains and hence are expected to substantially improve the accuracy of our method. A set of 400 human proteins whose localizations are predicted by our new method will be experimentally tested in normal and cancer cell lines of human, using GFP-fusion and expression followed by visualization under confocal microscope. This step would allow us to determine the prediction accuracy of our method at each score threshold for each organelle. Using optimal score thresholds, proteome-wide predictions will be carried out and detailed catalogs of experimentally-known and predicted subcellular and suborganellar proteomes will be generated for all sequenced genomes of animal species. Additionally, a standalone software package for the improved method will be developed and released to the research community under the General Public License (GPL). An online web server will be developed to make predictions online, and to enable access to the cataloged data and to the software produced in this project. In summary, the proposed comprehensive system will deliver a 'gold-standard' dataset of experimentally established localizations, a novel methodology for prediction, experimental validation of predicted localizations, and a public web server to predict or to access datasets and the software tool developed in this project. These resources will prove to be very valuable to the biomedical research community in advancing the many facets of systems biology research.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Biomedical Informatics, Bioinformatics, and Cyberinfrastructure Enhancement Core
-
批准号:10478974
-
项目类别:
-
资助金额:$35.06万
-
财政年份:2016
-
负责人:CHITTIBABU GUDA
-
依托单位:
Biomedical Informatics, Bioinformatics, and Cyberinfrastructure Enhancement Core
-
批准号:10281662
-
项目类别:
-
资助金额:$35.35万
-
财政年份:2016
-
负责人:CHITTIBABU GUDA
-
依托单位:
Cataloging the subcellular and suborganellar proteomes of sequenced genomes
-
批准号:8331447
-
项目类别:
-
资助金额:$21.83万
-
财政年份:2009
-
负责人:CHITTIBABU GUDA
-
依托单位:
Cataloging the subcellular and suborganellar proteomes of sequenced genomes
-
批准号:8234952
-
项目类别:
-
资助金额:$21.83万
-
财政年份:2009
-
负责人:CHITTIBABU GUDA
-
依托单位:
Cataloging the subcellular and suborganellar proteomes of sequenced genomes
-
批准号:8538440
-
项目类别:
-
资助金额:$17.56万
-
财政年份:2009
-
负责人:CHITTIBABU GUDA
-
依托单位:
Core C: Bioinformatics and Genomics Core
-
批准号:10627091
-
项目类别:
-
资助金额:$32.34万
-
财政年份:2009
-
负责人:CHITTIBABU GUDA
-
依托单位:
Core C - Bioinformatics and Genomics Core
-
批准号:9280513
-
项目类别:
-
资助金额:$12.09万
-
财政年份:2009
-
负责人:CHITTIBABU GUDA
-
依托单位:
Cataloging the subcellular and suborganellar proteomes of sequenced genomes
-
批准号:8138592
-
项目类别:
-
资助金额:$14.7万
-
财政年份:2009
-
负责人:CHITTIBABU GUDA
-
依托单位:
An integrated approach to infer and validate domain-domain interactions in protei
-
批准号:7367241
-
项目类别:
-
资助金额:$22.72万
-
财政年份:2008
-
负责人:CHITTIBABU GUDA
-
依托单位:
Nebraska Research Network in Functional Genomics
-
批准号:10624375
-
项目类别:
-
资助金额:$35.87万
-
财政年份:2001
-
负责人:CHITTIBABU GUDA
-
依托单位:
Nebraska Research Network in Functional Genomics
-
批准号:10426059
-
项目类别:
-
资助金额:$35.87万
-
财政年份:2001
-
负责人:CHITTIBABU GUDA
-
依托单位:
Bioinformatics Shared Resource
-
批准号:10491798
-
项目类别:
-
资助金额:$8.29万
-
财政年份:1997
-
负责人:CHITTIBABU GUDA
-
依托单位:
Bioinformatics (BISR)
-
批准号:9981645
-
项目类别:
-
资助金额:$10.33万
-
财政年份:1997
-
负责人:CHITTIBABU GUDA
-
依托单位:
Bioinformatics Shared Resource
-
批准号:10270913
-
项目类别:
-
资助金额:$8.27万
-
财政年份:1997
-
负责人:CHITTIBABU GUDA
-
依托单位:
Core C - Bioinformatics and Genomics Core
-
批准号:9755298
-
项目类别:
-
资助金额:$11.93万
-
财政年份:--
-
负责人:CHITTIBABU GUDA
-
依托单位:
UNMC Bioinformatics Core
-
批准号:8899797
-
项目类别:
-
资助金额:$35.54万
-
财政年份:--
-
负责人:CHITTIBABU GUDA
-
依托单位:
UNMC Bioinformatics Core
-
批准号:9095404
-
项目类别:
-
资助金额:$34.74万
-
财政年份:--
-
负责人:CHITTIBABU GUDA
-
依托单位:
Bioinformatics (BISR)
-
批准号:9755217
-
项目类别:
-
资助金额:$10.33万
-
财政年份:--
-
负责人:CHITTIBABU GUDA
-
依托单位:
Nebraska Research Network in Functional Genomics
-
批准号:9900469
-
项目类别:
-
资助金额:$35.87万
-
财政年份:--
-
负责人:CHITTIBABU GUDA
-
依托单位:
海外基金