课题基金 / 基金详情

Standards-compliant software tools for curation and public deposition of proteomics data

Standards-compliant software tools for curation and public deposition of proteomics data
用于蛋白质组数据管理和公开存储的符合标准的软件工具
批准号:
BB/H024654/1
负责人:
Andrew Jones
金额:
$13.77万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2010
资助国家:
英国
项目状态:
已结题
起止时间:
2010 至 --

项目摘要

项目成果

Andrew Jones的其他基金

相似基金

相关文献

中文摘要
翻译
我们现在正处于生物学研究的“后基因组”时代,因为许多生物体的完整基因组成现在已经知道了。基因组序列本质上为我们提供了完整生物体的部件列表和描述生物体如何运作的蓝图。然而,基因组序列只是一个静态的表示-它在生物体的所有细胞中都是相同的,并且在生物体的一生中不会改变。所有的细胞功能都是由蛋白质提供的,每一种蛋白质都由基因组中的一个基因编码。在每个细胞中表达的一组蛋白质随着生物体的发育和置于不同的环境条件下(如细胞上的外部应力)而发生显着变化。研究人员感兴趣的是研究样品中的全套蛋白质,这被称为蛋白质组。世界各地的许多研究实验室现在正在进行这些蛋白质组学研究,使用质谱法同时识别大量蛋白质。蛋白质组学研究可以简单地用于鉴定感兴趣的样品中存在的蛋白质,或者将一种条件下存在的蛋白质与另一种条件下存在的蛋白质进行比较。这些技术可用于找到打开或关闭的蛋白质,例如在疾病过程中或感染寄生生物后,帮助我们了解反应的分子基础。这些研究可以产生大量的数据集,这些数据集首先由质谱仪产生,然后由选择用于数据分析的软件包产生。研究人员很难访问和分析其他实验室产生的数据,除非他们在实验室使用相同的软件包,因为每个软件通常只能使用自己的数据格式。一个全球性的联盟已经形成,以评估这些研究的数据是如何报告的,在这种情况下,我的团队领导了一项倡议,创建一个单一的数据格式来存储蛋白质组学研究的结果。我们最近发布了一种标准格式,用于从质谱数据中进行蛋白质鉴定,称为mzIdentML。在这个应用程序中,我们的目标是构建软件,以便世界各地的软件开发人员可以轻松构建新的应用程序或修改现有的应用程序,以便他们可以以这种格式读取和写入数据。这将意味着,可以使用各种软件包,而不仅仅是源实验室使用的软件,对一个实验室产生的数据进行分析。研究人员在发表研究报告时,也越来越多地将数据存储在公共数据库中。这将允许其他研究人员下载数据以检查他们的发现,并将已发布的数据与他们自己的结果相结合。如果有一个世界公认的标准格式,这些过程将大大简化。在此应用中,我们正在开发一套软件工具,用于对识别数据执行标准任务,例如执行统计分析以确定特定结果的重要性。这些工具都是开源的,因为我们打算让学术界和工业界的其他软件开发人员在他们的应用程序中重用这些工具。我们还在开发一个图形查看器,以便科学家可以以不同的方式可视化他们的数据,以帮助理解大型数据集中包含的结果。该浏览器将被集成到一个主要公共数据库的应用程序中,使科学家能够简单地分析他们的数据,然后在发表期刊文章时将其直接上传到公共数据库。这些进展将帮助从事蛋白质组学研究的科学家分享他们的数据。反过来,这将有助于蛋白质组学数据库的快速增长,这将有利于所有分子生物学研究人员,因为他们将有机会获得巨大的蛋白质数据集进行数据挖掘,从而实现新的生物学发现。
英文摘要
We are now in the 'post-genomic' era of biological research since the complete genetic makeup of many organisms is now known. The genome sequence essentially provides us with the parts list of the complete organism and the blueprint describing how the organism functions. The genome sequence is only a static representation though - it is the same in all cells of an organism and does not change during the organism's lifetime. All cellular functions are provided by proteins, which are each encoded by a gene within the genome. The set of proteins expressed in each cell changes dramatically as the organism develops and as it is placed under different environmental conditions, such as external stresses on cells. Researchers are interested in studying the complete set of proteins in a sample, which is called the proteome. Many research laboratories around the world are now performing these proteomics studies, using mass spectrometry to identify large sets of proteins concurrently. Proteomics studies may be used simply to identify the proteins present in the samples of interest, or to compare the proteins present in one condition compared with another. These techniques can be used to find the proteins that are switched on or off, for example during a disease process or following infection with a parasitic organism, helping us to understand the molecular basis of the response. These studies can generate vast data sets, which are produced first by the mass spectrometer and then by the software package chosen for data analysis. It can be difficult for researchers to access and analyse data produced in other laboratories, unless they use the same software package in their laboratory, because each piece of software generally only works with its own data format. A worldwide consortium has formed to standardise how data from these studies are reported, and in this context, my group has led an initiative to create a single data format for storing the results of proteomics studies. We have recently released a standard format for protein identifications made from mass spectrometry data, called mzIdentML. In this application, we are aiming to build software so that it is easy for software developers around the world to build new applications or alter existing applications so that they can read and write data in this format. This will mean that data produced in one laboratory could be analysed using a variety of software packages, not just the software in use in the source laboratory. There is also a growing movement for researchers to store their data in public databases when they publish a study. This will allow other researchers to download the data to check their findings and to integrate the published data with their own results. These processes will be greatly simplified by having a worldwide-accepted standard format. In this application, we are developing a set of software tools for performing standard tasks on identification data, for example performing statistical analysis to determine significance of particular results. These tools are open-source, as we intend for other software developers, working both in academia and industry, to re-use these tools in their applications. We are also developing a graphical viewer so that scientists can visualise their data in different ways to help understand the results contained in the large data sets. The viewer will be integrated into an application produced by one of the main public databases, making it simple for scientists to analyse their data and then upload it directly to the public database when they publish a journal article. These developments will help scientists working in proteomics to share their data. In turn, this will help proteomics databases to grow rapidly, which will benefit all molecular biology researchers as they will have access to huge protein data sets for data mining, allowing new biological discoveries to be made.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1074/mcp.o113.036681
发表时间: 2014-10
期刊: Molecular & cellular proteomics : MCP
影响因子: --
作者: [Griss J, Jones AR, Sachsenberg T, Walzer M, Gatto L, Hartler J, Thallinger GG, Salek RM, Steinbeck C, Neuhauser N, Cox J, Neumann S, Fan J, Reisinger F, Xu QW, Del Toro N, Pérez-Riverol Y, Ghali F, Bandeira N, Xenarios I, Kohlbacher O, Vizcaíno JA, Hermjakob H]
通讯作者: Hermjakob H
DOI: 10.1002/pmic.201400553
发表时间: 2015-08
期刊: Proteomics
影响因子: 3.4
作者: [Krishna R, Xia D, Sanderson S, Shanmugasundram A, Vermont S, Bernal A, Daniel-Naguib G, Ghali F, Brunk BP, Roos DS, Wastling JM, Jones AR]
通讯作者: Jones AR
DOI: 10.1093/database/bat009
发表时间: 2013
期刊: Database : the journal of biological databases and curation
影响因子: --
作者: [Mayer G, Montecchi-Palazzi L, Ovelleiro D, Jones AR, Binz PA, Deutsch EW, Chambers M, Kallhardt M, Levander F, Shofstahl J, Orchard S, Vizcaíno JA, Hermjakob H, Stephan C, Meyer HE, Eisenacher M, HUPO-PSI Group]
通讯作者: HUPO-PSI Group
Controlled vocabularies and ontologies in proteomics: overview, principles and practice.
蛋白质组学中的受控词汇和本体论:概述,原理和实践。
DOI: 10.1016/j.bbapap.2013.02.017
发表时间: 2014-01
期刊: Biochimica et biophysica acta
影响因子: --
作者: [Mayer G, Jones AR, Binz PA, Deutsch EW, Orchard S, Montecchi-Palazzi L, Vizcaíno JA, Hermjakob H, Oveillero D, Julian R, Stephan C, Meyer HE, Eisenacher M]
通讯作者: Eisenacher M
共 6 条
    BBSRC-NSF/BIO. Globally harmonized re-analysis of Data Independent Acquisition (DIA) proteomics datasets enables the creation of new resources
    • 批准号:
      BB/X002020/1
    • 项目类别:
      Research Grant
    • 资助金额:
      $43.29万
    • 财政年份:
      2023
    • 负责人:
      Andrew Jones
    • 依托单位:
    BBSRC-NSF/BIO PanOryza: Globally coordinated genomes, proteomes and pathways for rice
    • 批准号:
      BB/T015691/1
    • 项目类别:
      Research Grant
    • 资助金额:
      $64.13万
    • 财政年份:
      2020
    • 负责人:
      Andrew Jones
    • 依托单位:
    GRAPPA - Global compRehensive Atlas of Peptide and Protein Abundance
    • 批准号:
      BB/T019557/1
    • 项目类别:
      Research Grant
    • 资助金额:
      $41.84万
    • 财政年份:
      2020
    • 负责人:
      Andrew Jones
    • 依托单位:
    BBSRC-NSF/BIO PTMeXchange: Globally harmonized re-analysis and sharing of data on post-translational modifications
    • 批准号:
      BB/S017054/1
    • 项目类别:
      Research Grant
    • 资助金额:
      $39.56万
    • 财政年份:
      2019
    • 负责人:
      Andrew Jones
    • 依托单位:
    国内基金
    海外基金
    基于约束行为的柔性精微机构设计方法研究
    • 批准号:
      50975007
    • 项目类别:
      面上项目
    • 资助金额:
      38.0万元
    • 批准年份:
      2009
    • 负责人:
      毕树生
    • 依托单位: