Galaxy Workflows for Proteomics Informed by Transcriptomics (PIT)
Galaxy Workflows for Proteomics Informed by Transcriptomics (PIT)
批准号:
BB/K016075/1
负责人:
Conrad Bessant
金额:
$13.82万
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2013
资助国家:
英国
项目状态:
已结题
起止时间:
2013 至 --
中文摘要
确定哪些蛋白质存在于给定的生物样品中,以及以何种数量存在,对于理解许多生物过程至关重要。一种称为“鸟枪蛋白质组学”的技术已成为解决这一问题的首选方法。在鸟枪法蛋白质组学分析中,首先使用裂解酶将蛋白质分解成更容易分析的片段(肽),然后使用液相色谱法(LC)分离,然后单独注射到串联质谱仪(MS/MS)中,将肽分解成片段,产生产物离子的光谱,可以被认为是每个肽的指纹。软件用于将获得的光谱与肽相匹配,然后这些肽鉴定用于推断蛋白质的存在。确定每个获得的光谱代表哪种肽显然是鸟枪法蛋白质组学的关键部分。从理论上讲,因为我们了解了肽片段化的原理,所以应该可以获得任何肽谱,并计算出它所来自的肽的序列。在实践中,这通常是太困难了,因为不完美的MS/MS光谱和可能存在的大量肽的组合使得不正确的鉴定非常可能。为了避免这个问题,蛋白质鉴定软件试图将肽谱仅与那些可能合理预期在样品中的肽序列相匹配。目前,这是通过搜索已知被研究物种产生的所有蛋白质(“蛋白质组”)的序列来完成的,这些蛋白质从在线数据库(例如UniProt)下载。然而,高质量的蛋白质组仅适用于少数物种。如果你想对一个没有蛋白质组的物种的样本进行蛋白质组学研究,或者对一个涉及多个物种或未知物种的实验样本进行蛋白质组学研究,该怎么办?我们最近开发了(并测试和发表)解决这个问题的方法,我们称之为转录组学(PIT)的蛋白质组学。PIT的关键是创建可能存在的蛋白质的样本特定列表,这些蛋白质来源于样本中发现的基因转录本。转录本是用于制造蛋白质的基因的拷贝,因此通过了解样本中存在哪些转录本,我们可以预测可能存在哪些蛋白质。转录本是通过使用称为RNA-seq的下一代测序技术发现的。直到最近,RNA-seq还涉及将短读段映射到参考基因组,但现在已有可以从头组装转录本的软件。因此,PIT方法使得在参考蛋白质组(或基因组)不可用时识别和量化复杂样本中的蛋白质成为可能。这为没有很好注释基因组的物种(包括许多害虫,病原体和植物)开辟了许多新的研究领域,也为来自多个物种的蛋白质存在的实验(所谓的“元蛋白质组学”)或蛋白质组发生变化的实验(例如在病毒感染期间)开辟了新的研究领域。PIT还有许多额外的附带好处,例如能够找到研究个体特有的蛋白质变体(即不存在于任何参考蛋白质组中),以及注释基因组的可能性。目前,PIT方法的主要挑战是整合转录组和蛋白质组数据所需的数据分析的复杂性,并以对生物学家有用的方式报告结果。因此,本提案的目的是将一套易于使用的连接软件工具放在一起,使典型的实验室科学家能够在可接受的时间范围内进行必要的数据分析,而无需生物信息学支持。为了帮助实现这一目标,我们计划在流行的Galaxy框架内实施该软件。Galaxy提供了一个易于使用的Web浏览器界面,并可以利用强大的计算资源。
英文摘要
Identifying which proteins are present in a given biological sample, and in what quantities, is essential to understanding many biological processes. A technique called "shotgun proteomics" has become the method of choice for tackling this problem. In a shotgun proteomics analysis proteins are first broken down into more easily analysable segments (peptides) using a cleavage enzyme, then separated using liquid chromatography (LC), prior to individual injection into a tandem mass spectrometer (MS/MS), which breaks peptides into fragments, producing a spectrum of product ions that can be considered as a fingerprint for each peptide. Software is used to match the acquired spectra to peptides and these peptide identifications are then used to infer the presence of proteins. Working out which peptide is represented by each of the acquired spectra is clearly a crucial part of shotgun proteomics. In theory, because we understand the principles of peptide fragmentation, it should be possible to take any peptide spectrum and work out the sequence of the peptide from which it came. In practice this is usually too difficult because the combination of imperfect MS/MS spectra and the huge number of peptides that could potentially exist make incorrect identifications very likely. To circumvent this problem, protein identification software seeks to match peptide spectra only to those peptide sequences that might reasonably be expected to be in the sample. Currently this is done by searching against the sequences of all proteins that the species under study is known to produce (the "proteome"), downloaded from an online database (e.g. UniProt). However, high quality proteomes are only available for a small number of species. What if you want to do proteomics on a sample from a species for which a proteome is not available, or on a sample from an experiment involving multiple species, or unknown species?We recently developed (and tested, and published) a solution to this problem, which we call proteomics informed by transcriptomics (PIT). The key to PIT is the creation of a sample-specific list of proteins that may be present, derived from gene transcripts found in the sample. Transcripts are copies of genes that are used to make proteins, so by knowing which transcripts are present in a sample we can predict which proteins might be present. The transcripts are found by using a next generation sequencing technique called RNA-seq. Until very recently, RNA-seq involved mapping short reads to a reference genome, but software is now available that can assemble transcripts de novo.The PIT approach therefore makes it possible to identify and quantify proteins in complex samples when a reference proteome (or genome) is not available. This opens many new areas of research for species that do not have well annotated genomes (which include many pests, pathogens and plants), and also for experiments where proteins from multiple species are present (so-called "metaproteomics") or where the proteome is changing (e.g. during viral infection). There are also a number of additional spin-off benefits such as the ability to find protein variants that are specific to the individual under study (i.e. not present in any reference proteome), and possibility to annotate genomes.Currently, the main challenge of the PIT approach is the complexity of the data analysis necessary to integrate the transcriptomic and proteomic data and report results in a way that is useful to biologists. The aim of this proposal is therefore to put together a suite of easy to use connected software tools that enable the typical bench scientist to perform the necessary data analysis within an acceptable timescale with no bioinformatics support. To help achieve this we plan to implement the software within the popular Galaxy framework. Galaxy provides an easy to use web browser interface and can take advantage of powerful computing resources.
期刊论文(5)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1074/mcp.o115.048777
发表时间:
2015-11
期刊:
Molecular & cellular proteomics : MCP
影响因子:
--
作者:
[Fan J, Saha S, Barker G, Heesom KJ, Ghali F, Jones AR, Matthews DA, Bessant C]
通讯作者:
Bessant C
DOI:
10.1080/2159256x.2017.1362494
发表时间:
2017
期刊:
Mobile genetic elements
影响因子:
--
作者:
[Davidson AD, Matthews DA, Maringer K]
通讯作者:
Maringer K
Proteomics informed by transcriptomics for characterising active transposable elements and genome annotation in Aedes aegypti.
蛋白质组学通过转录组学告知,以表征伊蚊中的主动转座元件和基因组注释。
DOI:
10.1186/s12864-016-3432-5
发表时间:
2017-01-19
期刊:
BMC genomics
影响因子:
4.4
作者:
[Maringer K, Yousuf A, Heesom KJ, Fan J, Lee D, Fernandez-Sesma A, Bessant C, Matthews DA, Davidson AD]
通讯作者:
Davidson AD
DOI:
10.1093/nar/gkx906
发表时间:
2018-01-04
期刊:
Nucleic acids research
影响因子:
14.9
作者:
[Saha S, Chatzimichali EA, Matthews DA, Bessant C]
通讯作者:
Bessant C
DOI:
10.1093/nar/gky295
发表时间:
2018-06-01
期刊:
Nucleic acids research
影响因子:
14.9
作者:
[Saha S, Matthews DA, Bessant C]
通讯作者:
Bessant C
PIT-DB: A Resource for Sharing, Annotating and Analysing Translated Genomic Elements
-
批准号:BB/M020118/1
-
项目类别:Research Grant
-
资助金额:$15.64万
-
财政年份:2015
-
负责人:Conrad Bessant
-
依托单位:
Proteomics Goes Viral: Novel Resources for Identification and Quantification of Virus Proteins
-
批准号:BB/L018438/1
-
项目类别:Research Grant
-
资助金额:$18.9万
-
财政年份:2014
-
负责人:Conrad Bessant
-
依托单位:
An Integrated Open Source Software Resource for Quantitative Proteomics
-
批准号:BB/I001131/2
-
项目类别:Research Grant
-
资助金额:$0.93万
-
财政年份:2013
-
负责人:Conrad Bessant
-
依托单位:
An Integrated Open Source Software Resource for Quantitative Proteomics
-
批准号:BB/I001131/1
-
项目类别:Research Grant
-
资助金额:$26.63万
-
财政年份:2010
-
负责人:Conrad Bessant
-
依托单位:
X-tracker: a generic quantitation tool for MS-based proteomics:
-
批准号:BB/F016107/1
-
项目类别:Research Grant
-
资助金额:$13.14万
-
财政年份:2008
-
负责人:Conrad Bessant
-
依托单位:
Further Development of the Genome Annotating Proteomic Pipeline
-
批准号:BB/E01237X/1
-
项目类别:Research Grant
-
资助金额:$11.76万
-
财政年份:2007
-
负责人:Conrad Bessant
-
依托单位:
Bioinformatics for High Throughput Proteomics (Short Course)
-
批准号:BB/D007216/1
-
项目类别:Research Grant
-
资助金额:$6.82万
-
财政年份:2006
-
负责人:Conrad Bessant
-
依托单位:
海外基金