Transferable retention time prediction for Liquid Chromatography-Mass Spectrometry-based metabolomics
Transferable retention time prediction for Liquid Chromatography-Mass Spectrometry-based metabolomics
批准号:
425789784
负责人:
Professor Dr. Sebastian Böcker
金额:
$0.0万
依托单位:
依托单位国家:
德国
项目类别:
Research Grants
财政年份:
2019
资助国家:
德国
项目状态:
已结题
起止时间:
2018-12-31 至 2022-12-31
中文摘要
代谢物鉴定仍然是代谢组学的主要瓶颈。液质联用(LC-MS)是目前非靶向代谢组学中最常用的分析技术。目前,在一个典型的非靶向实验中,只有不到10%的光谱可以被注释。因此,迫切需要改进的代谢物鉴定工具。虽然单靠质量不能识别分子,但串联质谱会产生裂解光谱,可用于结构鉴定。最近,在电子计算机中,方法已经被开发并越来越多地被代谢组学社区使用,允许在诸如PubChem和ChemSpider的分子结构数据库中进行搜索。这样的结构数据库比任何波谱库都大许多个数量级,因此对分子结构的覆盖范围要广得多。但即使用串联质谱仪进行鉴定,也会导致大量的虚假鉴定。为了提高鉴定质量,必须报告两个独立的参数,如化学标准物质的质量和保留时间。今天,保留时间主要用于鉴定管道的后期阶段,主要基于与化学参比标准的比较。然而,如果在早期阶段使用保留时间显然是有益的,特别是在电子计算方法中;在这里,我们可以筛选候选对象,或者更好地,基于预测和观察保留时间的比较来修改候选对象的分数。本项目旨在更好地利用保留时间在基于LC-MS的非靶向代谢组学中识别小生物分子,使用可转移保留时间预测。预测将基于两步走的方法。首先,机器学习将被用来预测给定分子结构的保留序号;培训将基于从公共可用数据集中广泛精选的保留时间数据收集,以及参考代谢物标准的系统内部测量。与其质量不同,保留时间不是代谢物的特征,而是代谢物、固定相和流动相的组合。因此,除了代谢物的分子指纹外,我们还将使用所采用的色谱系统的特性来进行机器学习。在第二步中,保留序号将被映射到保留时间,使用已知和识别的物质作为映射的锚。保留顺序和保留时间预测将用于过滤假阳性反应对,并应用于线虫次生代谢的独立生物数据集。所有精选和获得的数据,用于预测保留顺序和保留时间的开源软件将免费提供给代谢组学社区。最后,将保留时间预测整合到CSI:FingerID评分中,以提高其代谢物识别率。
英文摘要
Metabolite identification still represent the major bottleneck in metabolomics. Liquid Chromatography-Mass Spectrometry (LC-MS) is the currently most employed analytical technique in untargeted metabolomics. Currently, less than 10% of spectra in a typical untargeted experiment can be annotated. Therefore, there is a strong need for improved tools for metabolite identification. While mass alone cannot identify molecules, tandem MS yields fragmentation spectra which can be used for structural elucidation. Recently, in silico approaches have been developed and are increasingly used by the metabolomics community, that allow to search in molecular structure databases such as PubChem and ChemSpider. Such structure databases are many orders of magnitude larger than any spectral library and, hence, have a much wider coverage of molecular structures. But even identification by tandem MS will result in numerous spurious identifications. To improve identification quality, two independent parameters, e.g. mass and retention time of a chemical reference standard have to be reported. Today, retention time is mainly used at a later stage of the identification pipeline, and mainly based on comparison with chemical reference standards. However, it would clearly be beneficial if retention times were used at an early stage, in particular for in silico methods; here, we could filter candidates or, even better, modify the scores of candidates based on comparing predicted and observed retention times.This project aims to make better use of retention times for the identification of small biomolecules in LC-MS based untargeted metabolomics, using transferable retention time prediction. Prediction will be based on a two-step approach. First, Machine Learning will be used to predict retention order numbers for give molecular structures; training will be based on an extensively curated collection of retention time data from public available datasets, as well as systematic in-house measurements for reference metabolite standards. In contrast to its mass, retention time is not a feature of a metabolite, but of the combination of metabolite, stationary and mobile phase. Therefore, we will use properties of the employed chromatographic system in addition to molecular fingerprints of metabolites for machine learning. In the second step, retention order numbers will be mapped to retention times, using known and identified substances as anchors of the mapping. Retention order and retention time prediction will be used to filter false positive reaction pairs, and applied to an independent biological dataset from C. elegans secondary metabolism.All curated and acquired data, open-source software for prediction of retention order and retention times will be made freely available to the metabolomics community. Finally, retention time prediction will be integrated into the CSI:FingerID scoring in order to improve its metabolite identification rates.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Identifying the unknowns: towards structural elucidation of small molecules using mass spectrometry
-
批准号:242259350
-
项目类别:Research Grants
-
资助金额:$0.0万
-
财政年份:2013
-
负责人:Professor Dr. Sebastian Böcker
-
依托单位:
FlipCut Supertrees: Große und akkurate Phylogenien schneller bestimmen
-
批准号:211926079
-
项目类别:Research Grants
-
资助金额:$0.0万
-
财政年份:2012
-
负责人:Professor Dr. Sebastian Böcker
-
依托单位:
Algorithms for the Analysis of Approximate Gene Cluster (3AGC)
-
批准号:156864160
-
项目类别:Research Grants
-
资助金额:$0.0万
-
财政年份:2010
-
负责人:Professor Dr. Sebastian Böcker
-
依托单位:
Identifying the unknowns: towards structural elucidation of small molecules using mass spectrometry
-
批准号:164582891
-
项目类别:Research Grants
-
资助金额:$0.0万
-
财政年份:2010
-
负责人:Professor Dr. Sebastian Böcker
-
依托单位:
Parameterized Algorithmics for Bioinformatics
-
批准号:162571619
-
项目类别:Research Grants
-
资助金额:$0.0万
-
财政年份:2009
-
负责人:Professor Dr. Sebastian Böcker
-
依托单位:
Informatische Methoden für Massenspektrometrie in der Genomik
-
批准号:5400926
-
项目类别:Independent Junior Research Groups
-
资助金额:$0.0万
-
财政年份:2003
-
负责人:Professor Dr. Sebastian Böcker
-
依托单位:
Project Harvester: Improving molecular fingerprint prediction through self-training
-
批准号:518231245
-
项目类别:Research Grants
-
资助金额:$0.0万
-
财政年份:--
-
负责人:Professor Dr. Sebastian Böcker
-
依托单位:
Identifying the Unknowns: Fragmentation Trees and Molecular Fingerprints
-
批准号:324792648
-
项目类别:Research Grants
-
资助金额:$0.0万
-
财政年份:--
-
负责人:Professor Dr. Sebastian Böcker
-
依托单位:
海外基金