Problems, principles and progress in computational annotation of NMR metabolomics data.

Problems, principles and progress in computational annotation of NMR metabolomics data.
复制标题

DOI:
10.1007/s11306-022-01962-z
复制
发表时间:
2022-12-05
期刊:
Metabolomics : Official journal of the Metabolomic Society
影响因子:
--
通讯作者:
--
中科院分区:
其他
文献类型:
--
作者:

文献摘要

参考文献

被引文献

相似文献

化合物鉴定仍然是核磁共振代谢组学数据开发过程中的一个关键瓶颈,尤其是1H-1维数据。随着参考化合物光谱数据库的增长,工作流程已演变为在很大程度上依赖其搜索功能,通过生成在复杂混合物数据中发现的潜在代谢物列表来促进这一过程,便于注释和鉴定。然而,验证和传达注释的方法通常由专家知识指导,因此尽管反复努力调整实践和定义社区标准,但方法的变化性很高。这篇综述旨在通过讨论光谱匹配的关键思想来扩大自动标注工具的应用,并开始描述一组术语来对这些信息进行分类,从而提高传递标注置信度的标准。此外,我们希望这次审查将促进化学数据科学家、软件开发人员和核磁共振代谢组学社区之间日益增长的合作,以帮助开发长期的软件解决方案。我们首先简要讨论典型的非靶向核磁共振识别工作流程。我们区分了注释(假设生成、过滤)和识别(假设测试、验证),并注意到不同的核磁共振数据特征对注释的效用。然后我们涉及到注释的三个部分:(1)查询的生成,(2)将查询与参考数据匹配,以及(3)用于验证的潜在匹配的评分和置信度估计。在这样做的过程中,我们从它们所利用的结构信息的角度强调了自动化和半自动注释的现有方法,以及如何用计算来表示这些信息。
Compound identification remains a critical bottleneck in the process of exploiting Nuclear Magnetic Resonance (NMR) metabolomics data, especially for 1H 1-dimensional (1H 1D) data. As databases of reference compound spectra have grown, workflows have evolved to rely heavily on their search functions to facilitate this process by generating lists of potential metabolites found in complex mixture data, facilitating annotation and identification. However, approaches for validating and communicating annotations are most often guided by expert knowledge, and therefore are highly variable despite repeated efforts to align practices and define community standards. This review is aimed at broadening the application of automated annotation tools by discussing the key ideas of spectral matching and beginning to describe a set of terms to classify this information, thus advancing standards for communicating annotation confidence. Additionally, we hope that this review will facilitate the growing collaboration between chemical data scientists, software developers and the NMR metabolomics community aiding development of long-term software solutions. We begin with a brief discussion of the typical untargeted NMR identification workflow. We differentiate between annotation (hypothesis generation, filtering), and identification (hypothesis testing, verification), and note the utility of different NMR data features for annotation. We then touch on three parts of annotation: (1) generation of queries, (2) matching queries to reference data, and (3) scoring and confidence estimation of potential matches for verification. In doing so, we highlight existing approaches to automated and semi-automated annotation from the perspective of the structural information they utilize, as well as how this information can be represented computationally.
DOI: 10.1039/d1np00023c
发表时间: 2021-11-17
影响因子: 11.9
作者:
Beniddir MA;Kang KB;Genta-Jouve G;Huber F;Rogers S;van der Hooft JJJ
通讯作者: van der Hooft JJJ
TOCCATA:定制的碳总相关光谱NMR代谢组学数据库。
DOI: 10.1021/ac302197e
发表时间: 2012-11-06
影响因子: 7.4
作者:
Bingol, Kerem;Zhang, Fengli;Bruschweiler-Li, Lei;Brueschweiler, Rafael
通讯作者: Brueschweiler, Rafael
DOI: 10.1016/j.csbj.2016.02.005
发表时间: 2016
影响因子: 6
作者:
Dona AC;Kyriakides M;Scott F;Shephard EA;Varshavi D;Veselkov K;Everett JR
通讯作者: Everett JR
DOI: 10.1162/netn_a_00199
发表时间: 2021
期刊: Network neuroscience (Cambridge, Mass.)
影响因子: --
作者:
Frigo M;Cruciani E;Coudert D;Deriche R;Natale E;Deslauriers-Gauthier S
通讯作者: Deslauriers-Gauthier S
DOI: 10.1021/acs.jproteome.9b00872
发表时间: 2020-08-07
影响因子: 4.4
作者:
Charris-Molina, Andres;Riquelme, Gabriel;Hoijemberg, Pablo A.
通讯作者: Hoijemberg, Pablo A.