Problems, principles and progress in computational annotation of NMR metabolomics data.
Problems, principles and progress in computational annotation of NMR metabolomics data.
复制标题
DOI:
10.1007/s11306-022-01962-z
复制
发表时间:
2022-12-05
期刊:
影响因子:
--
通讯作者:
中科院分区:
文献类型:
--
作者:
Compound identification remains a critical bottleneck in the process of exploiting Nuclear Magnetic Resonance (NMR) metabolomics data, especially for 1H 1-dimensional (1H 1D) data. As databases of reference compound spectra have grown, workflows have evolved to rely heavily on their search functions to facilitate this process by generating lists of potential metabolites found in complex mixture data, facilitating annotation and identification. However, approaches for validating and communicating annotations are most often guided by expert knowledge, and therefore are highly variable despite repeated efforts to align practices and define community standards. This review is aimed at broadening the application of automated annotation tools by discussing the key ideas of spectral matching and beginning to describe a set of terms to classify this information, thus advancing standards for communicating annotation confidence. Additionally, we hope that this review will facilitate the growing collaboration between chemical data scientists, software developers and the NMR metabolomics community aiding development of long-term software solutions. We begin with a brief discussion of the typical untargeted NMR identification workflow. We differentiate between annotation (hypothesis generation, filtering), and identification (hypothesis testing, verification), and note the utility of different NMR data features for annotation. We then touch on three parts of annotation: (1) generation of queries, (2) matching queries to reference data, and (3) scoring and confidence estimation of potential matches for verification. In doing so, we highlight existing approaches to automated and semi-automated annotation from the perspective of the structural information they utilize, as well as how this information can be represented computationally.
登录
查看更多内容
影响因子:
11.9
作者:
Beniddir MA;Kang KB;Genta-Jouve G;Huber F;Rogers S;van der Hooft JJJ
通讯作者:
van der Hooft JJJ
影响因子:
7.4
作者:
Bingol, Kerem;Zhang, Fengli;Bruschweiler-Li, Lei;Brueschweiler, Rafael
通讯作者:
Brueschweiler, Rafael
影响因子:
6
作者:
Dona AC;Kyriakides M;Scott F;Shephard EA;Varshavi D;Veselkov K;Everett JR
通讯作者:
Everett JR
DOI:
10.1162/netn_a_00199
发表时间:
2021
期刊:
Network neuroscience (Cambridge, Mass.)
影响因子:
--
作者:
Frigo M;Cruciani E;Coudert D;Deriche R;Natale E;Deslauriers-Gauthier S
通讯作者:
Deslauriers-Gauthier S
影响因子:
4.4
作者:
Charris-Molina, Andres;Riquelme, Gabriel;Hoijemberg, Pablo A.
通讯作者:
Hoijemberg, Pablo A.