BBSRC-NSF/BIO. Globally harmonized re-analysis of Data Independent Acquisition (DIA) proteomics datasets enables the creation of new resources
BBSRC-NSF/BIO. Globally harmonized re-analysis of Data Independent Acquisition (DIA) proteomics datasets enables the creation of new resources
批准号:
BB/X001911/1
负责人:
Juan Antonio Vizcaino
金额:
$62.82万
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2023
资助国家:
英国
项目状态:
未结题
起止时间:
2023 至 --
中文摘要
蛋白质是重要的分子,它执行生物体每个细胞中发生的大多数活动,例如运输物质和提供结构支持。蛋白质组是在一定条件下在给定时间内系统或生物体中所有蛋白质的完整集合,蛋白质组学是对蛋白质组的大规模研究。蛋白质组学应用于生物学的许多方面,因为它可以告诉我们很多关于系统或生物体如何工作的信息,并可以提供有关疾病和潜在治疗的重要信息。蛋白质组学研究中使用的主要技术是质谱法(MS),其工作原理是将混合蛋白质样品分解成小片段,对其进行分类,然后报告其质量。该信息用于确定蛋白质的身份和量。最近,一种称为数据独立采集(DIA)的MS方法已经变得流行。传统的MS,称为数据相关采集(DDA),偏向于具有最强信号的片段,但DIA不受此限制。这意味着DIA允许研究人员量化即使是非常小数量的蛋白质,从而更好地表示蛋白质组。光谱库是用于DIA数据分析的预先注释的实验MS输出的集合。最近,光谱库已经使用机器学习开发,这为蛋白质组学研究的新人工智能(AI)方法提供了一个很好的机会。总的来说,定量DIA数据非常丰富,因为它代表了蛋白质组的全面数字记录,可以使用不同的工具和方法进行分析。参与该项目的团队一直致力于通过ProteomeXchange(PX)联盟在全球范围内免费提供DIA蛋白质组学数据,并确保通过蛋白质组学标准倡议(PSI)使用一致的标准生成和报告这些数据。这些公开的数据为研究人员重新确认原始结果并获得新的见解提供了很好的机会。然而,迄今为止,重新分析的努力非常有限。这可能是由于DIA数据分析的复杂性,也是因为缺乏可用的光谱库。我们的项目旨在通过重新分析DIA蛋白质组学数据集产生新的知识,并创建新的基础设施来更好地支持公共DIA蛋白质组学数据和光谱库。此外,我们将创建新的基础设施,使光谱库可查找,可解释,可互操作和可重用(FAIR),这将提高已发表研究的重现性。为了实现这些目标,我们将通过重新分析手动管理的公共DIA定量数据集来产生可靠和高质量的蛋白质表达(即蛋白质生产)和丰度信息,我们将通过PX和EMBL-EBI的Expression Atlas免费提供这些信息,供非蛋白质组学专家使用。我们还将使用DIA重新分析创建不同生物条件下的蛋白质共表达和丰度图,并通过PX提供。这将是第一次在如此大量的DIA蛋白质组学数据上生成这些图谱,并将利用DIA数据集的独特优势,如大小和覆盖范围。此外,我们将开发新的基础设施和数据标准,使DIA蛋白质组学数据,并作为一个关键点,光谱库公平。这将涉及创建开源工具和基础设施,并制定PSI标准。该项目将产生的共表达图谱,基础设施和标准将使广泛的生物和生物医学领域的研究人员受益,并将提供加强和连接现有研究成果的能力。我们将广泛宣传我们的工作,培训和协助研究人员充分利用这些宝贵的资源。
英文摘要
Proteins are important molecules that carry out most of the activities that take place in each cell of an organism, such as transporting substances and providing structural support. A proteome is the complete set of all the proteins in a system or organism under certain conditions at a given time, and proteomics is the large-scale study of proteomes. Proteomics applies to many parts of biology as it can tell us a lot about how a system or organism works, and can provide vital information about illnesses and potential treatments.The main technique used in proteomics research is mass spectrometry (MS), which works by breaking up a mixed protein sample into small fragments, sorting them and then reporting their mass. This information is used to determine the identity and amount of the proteins. Recently, a MS approach called data independent acquisition (DIA) has become popular. Traditional MS, called data dependent acquisition (DDA), is biased towards the fragments that have the strongest signal, but DIA is not limited by this. This means that DIA allows researchers to quantify proteins that are present even in very small numbers, allowing for better representation of the proteome. Spectral libraries are collections of pre-annotated experimental MS outputs that are used in DIA data analysis. Recently spectral libraries have been developed using machine learning, which provides a great opportunity for novel artificial intelligence (AI) approaches to proteomics research. Overall, quantitative DIA data is very rich, as it represents a comprehensive digital record of the proteome that can be analysed using different tools and approaches over time.The groups involved in this project have been working to make DIA proteomics data freely available worldwide via the ProteomeXchange (PX) consortium, and to ensure that this data is generated and reported using consistent standards via the Proteomics Standards Initiative (PSI). This publicly-available data provides a great opportunity for researchers to reconfirm original results and obtain new insights. However, there have so far been very limited re-analysis efforts. This may be due to the complex nature of DIA data analysis, and also because of a lack of availability of spectral libraries.Our project aims to address this by generating new knowledge coming from the re-analysis of DIA proteomics datasets and creating novel infrastructure to better support public DIA proteomics data and spectral libraries. Additionally, we will create novel infrastructure for making spectral libraries Findable, Accessible, Interoperable and Re-usable (FAIR), which will enhance the reproducibility of published studies. To achieve these goals we will produce reliable and high-quality protein expression (i.e. protein production) and abundance information from the re-analysis of manually curated public DIA quantitative datasets and we will make these freely available in PX and via EMBL-EBI's Expression Atlas, to be consumed by non-experts in proteomics. We will also create protein co-expression and abundance maps for different biological conditions using the DIA re-analyses and make them available via PX. This would be the first time that these maps are generated on such large amounts of DIA proteomics data and will take advantage of the unique advantages, such as size and coverage, of DIA datasets. Further, we will develop novel infrastructure and data standards to make DIA proteomics data and, as a key point, spectral libraries FAIR. This will involve creating open source tools and infrastructure, and developing PSI standards.The co-expression maps, infrastructure and standards that will be generated by this project will benefit researchers across a wide range of biological and biomedical fields, and will provide the ability to strengthen and connect existing research findings. We will disseminate our work widely to train and assist researchers in making full use of these valuable resources.
期刊论文(3)
专著(0)
科研奖励(0)
会议论文
DOI:
10.1093/nar/gkad1021
发表时间:
2024-01-05
期刊:
NUCLEIC ACIDS RESEARCH
影响因子:
14.9
作者:
[George, Nancy, Fexova, Silvie, Fuentes, Alfonso Munoz, Madrigal, Pedro, Bi, Yalan, Iqbal, Haider, Kumbham, Upendra, Nolte, Nadja Francesca, Zhao, Lingyun, Thanki, Anil S., Yu, Iris D., Marugan Calles, Jose C., Erdos, Karoly, Vilmovsky, Liora, Kurri, Sandeep R., Vathrakokoili-Pournara, Anna, Osumi-Sutherland, David, Prakash, Ananth, Wang, Shengbo, Tello-Ruiz, Marcela K., Kumari, Sunita, Ware, Doreen, Goutte-Gattat, Damien, Hu, Yanhui, Brown, Nick, Perrimon, Norbert, Vizcaino, Juan Antonio, Burdett, Tony, Teichmann, Sarah, Brazma, Alvis, Papatheodorou, Irene]
通讯作者:
Papatheodorou, Irene
The Open Data Exchange Ecosystem in Proteomics: Evolving its Utility
-
批准号:EP/Y035984/1
-
项目类别:Research Grant
-
资助金额:$16.81万
-
财政年份:2024
-
负责人:Juan Antonio Vizcaino
-
依托单位:
3D-Proteomics: FAIRification of proteomics data for comprehensive integration with structural biology information
-
批准号:BB/V018779/1
-
项目类别:Research Grant
-
资助金额:$89.39万
-
财政年份:2022
-
负责人:Juan Antonio Vizcaino
-
依托单位:
GRAPPA - Global compRehensive Atlas of Peptide and Protein Abundance
-
批准号:BB/T019670/1
-
项目类别:Research Grant
-
资助金额:$85.6万
-
财政年份:2021
-
负责人:Juan Antonio Vizcaino
-
依托单位:
BBSRC-NSF/BIO PTMeXchange: Globally harmonized re-analysis and sharing of data on post-translational modifications
-
批准号:BB/S01781X/1
-
项目类别:Research Grant
-
资助金额:$59.0万
-
财政年份:2019
-
负责人:Juan Antonio Vizcaino
-
依托单位:
In silico mass spectrometry for biologists: Tools and resources for next-generation proteomics
-
批准号:BB/P024599/1
-
项目类别:Research Grant
-
资助金额:$56.67万
-
财政年份:2017
-
负责人:Juan Antonio Vizcaino
-
依托单位:
国内基金
海外基金
登录
查看更多内容
SYNJ1蛋白片段通过促进突触蛋白NSF聚集在帕金森病发生中的机制研究
-
批准号:--
-
项目类别:青年科学基金项目
-
资助金额:30万元
-
批准年份:2022
-
负责人:邹利
-
依托单位:
NSF蛋白亚硝基化修饰所介导的GluA2 containing-AMPA受体膜稳定性在卒中后抑郁中的作用及机制研究
-
批准号:82071300
-
项目类别:面上项目
-
资助金额:55.0万元
-
批准年份:2020
-
负责人:方琪
-
依托单位:
参加中美(NSFC-NSF)生物多样性项目评审会
-
批准号:--
-
项目类别:国际(地区)合作与交流项目
-
资助金额:2万元
-
批准年份:2019
-
负责人:贺金生
-
依托单位:
参加中美(NSFC-NSF)生物多样性项目评审会
-
批准号:31981220281
-
项目类别:国际(地区)合作与交流项目
-
资助金额:2.3万元
-
批准年份:2019
-
负责人:张全发
-
依托单位:
中美(NSFC-NSF)EEID联合评审会
-
批准号:--
-
项目类别:国际(地区)合作与交流项目
-
资助金额:2.6万元
-
批准年份:2019
-
负责人:肖立华
-
依托单位:
中美(NSFC-NSF)EEID联合评审会
-
批准号:81981220037
-
项目类别:国际(地区)合作与交流项目
-
资助金额:2.1万元
-
批准年份:2019
-
负责人:段广才
-
依托单位:
中美(NSFC-NSF)EEID联合评审会
-
批准号:--
-
项目类别:国际(地区)合作与交流项目
-
资助金额:1.2万元
-
批准年份:2019
-
负责人:王四宝
-
依托单位:
Mon1b 协同NSF调控早期内吞体膜融合的机制研究
-
批准号:31671397
-
项目类别:面上项目
-
资助金额:67.0万元
-
批准年份:2016
-
负责人:李红昌
-
依托单位: