融入三代测序全长转录组信息的新蛋白鉴定方法
批准号:
32060149
项目类别:
地区科学基金项目
资助金额:
35.0 万元
负责人:
谢尚潜
依托单位:
学科分类:
生物数据资源与分析方法
结题年份:
2024
批准年份:
2020
项目状态:
已结题
项目参与者:
谢尚潜
中文摘要
新蛋白鉴定是蛋白质组学研究的重要内容,基于肽段从头测序的传统方法鉴定新蛋白准确率低,难以大规模应用。三代测序全长转录组具有全长转录本无需拼装以及大规模发现新转录本的优势,整合全长转录组信息和串联质谱技术为大规模鉴定新蛋白提供了可能。最新研究表明三代测序PacBio平台全长转录组序列信息优化蛋白搜索数据库可大规模鉴定新肽段,我们课题组前期研究表明二代测序转录组丰度信息能有效提高蛋白鉴定量。但目前利用三代测序全长转录组序列和丰度多维信息发现新蛋白的计算方法还有待进一步研究。因此本项目拟建立一套完整的融入三代测序全长转录组序列和丰度两维信息的新蛋白鉴定方法,新方法不仅将全长转录本序列信息用于构建蛋白搜索数据库以大规模发现新蛋白,而且将转录本丰度信息融入肽段匹配打分模型以提高蛋白鉴定量。新方法能有效提高蛋白鉴定量并大规模发现新蛋白,为蛋白质组学研究提供新的方法参考和技术支持。
英文摘要
The identification of new proteins is an important part of proteomics, but traditional proteomics uses peptide de novo sequencing to identify new proteins with low accuracy and is difficult to apply on a large scale. The full-length transcriptome from third-generation sequencing has the advantages of full-length transcripts without assembly and large-scale discovery of new transcripts. The integration of full-length transcriptome and tandem mass spectrometry provides the possibility for large-scale identification of new proteins. Recent studies have shown that using the full-length transcriptome sequence information of the third-generation sequencing PacBio platform to optimize the protein search database can be used to identify new peptides on a large scale. Furthermore, previous study of our group has shown that the transcriptome abundance information of RNA-seq from next-generation sequencing can effectively improve the amount of protein identification. However, the computational method of new proteins identification by using full-length transcriptional sequence and abundance information from third-generation sequencing remains to be further explored. Therefore, this project intends to develop an algorithm for new protein identification that integrates the full-length transcriptional sequence and abundance information from third-generation sequencing platforms. The new algorithm not only integrates the sequence information of full-length transcripts into the peptide search database to identify new proteins on a large scale, but also integrates transcript abundance information into the peptide matching scoring model to improve protein identification. The new strategy can effectively improve protein identification and discover new proteins on a large scale, providing a new method for reference and technical support for proteomics research.
新蛋白鉴定是蛋白质组学研究的重要内容,基于肽段从头测序的传统方法鉴定新蛋白通量低,难以大规模应用。三代测序全长转录组具有全长转录本无需拼装以及大规模发现新转录本的优势,整合全长转录组信息和串联质谱技术为大规模鉴定新蛋白提供了可能。三代测序PacBio和ONT平台的全长转录组序列信息优化蛋白搜索数据库可大规模鉴定新肽段,目前利用三代测序全长转录组序列和丰度多维信息发现新蛋白的计算方法还有待进一步研究。本项目利用三代测序全长转录本优势,建立了一套完整的融入三代测序全长转录组序列和丰度两维信息的新蛋白鉴定方法,主要内容包括:1.建立了全长转录本的识别方法;2.建立了全长转录本丰度FPKM与蛋白鉴定的定量化模型;3.建立了整合全长转录本FPKM的肽段打分模型;4.应用新方法肺癌和结肠癌肿瘤数据鉴定新蛋白,建立了新蛋白数据库NProDB。新方法不仅将全长转录本序列信息用于构建蛋白搜索数据库以大规模发现新蛋白,而且将转录本丰度信息融入肽段匹配打分模型以提高蛋白鉴定量。新方法能有效提高蛋白鉴定量并大规模发现新蛋白,为蛋白质组学研究提供新的方法参考和技术支持。
基于三代测序全长转录组的特异性Isoform识别方法研究及特征分析
-
批准号:31760316
-
项目类别:地区科学基金项目
-
资助金额:35.0万元
-
批准年份:2017
-
负责人:谢尚潜
-
依托单位:
基于多组学先验信息的串联质谱数据库搜索方法研究及应用
-
批准号:31600667
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2016
-
负责人:谢尚潜
-
依托单位:
国内基金
海外基金