Bayesian networks for mass spectrometric metabolite identification via molecular fingerprints.

Bayesian networks for mass spectrometric metabolite identification via molecular fingerprints.
复制标题

DOI:
10.1093/bioinformatics/bty245
复制
发表时间:
2018-07-01
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Böcker S
Böcker S
中科院分区:
其他
文献类型:
--
作者:
Ludwig M;Dührkop K;Böcker S

文献摘要

参考文献

被引文献

相似文献

代谢产物,参与细胞反应的小分子,提供细胞状态的直接功能特征。非靶向代谢组学实验通常依赖于串联质谱来识别生物样品中的数千种化合物。最近,我们提出了CSI:FingerID用于使用串联质谱数据在分子结构数据库中搜索。CSI:FingerID预测一个编码查询化合物结构的分子指纹,然后使用它来搜索分子结构数据库,如PubChem。假设构成指纹的分子性质之间的独立性,执行预测查询指纹和确定性目标指纹的评分。我们提出了一个评分,考虑到分子特性之间的依赖关系。和以前一样,我们使用机器学习来预测分子特性的后验概率。分子特性之间的相似性被建模为贝叶斯树网络;树结构是从实例数据动态估计的。对于每个边,我们还估计两个随机变量之间的期望协方差。对于固定的边际概率,我们使用已知的协方差估计条件概率。现在,可以计算每个候选人的校正后验概率,并根据该分数对候选人进行排名。建模依赖关系将CSI:FingerID的识别率提高了2.85个百分点。新的贝叶斯评分(固定树)已集成到SIRIUS 4.0中(https://bio.informatik.uni-jena.de/software/sirius/)。
Metabolites, small molecules that are involved in cellular reactions, provide a direct functional signature of cellular state. Untargeted metabolomics experiments usually rely on tandem mass spectrometry to identify the thousands of compounds in a biological sample. Recently, we presented CSI:FingerID for searching in molecular structure databases using tandem mass spectrometry data. CSI:FingerID predicts a molecular fingerprint that encodes the structure of the query compound, then uses this to search a molecular structure database such as PubChem. Scoring of the predicted query fingerprint and deterministic target fingerprints is carried out assuming independence between the molecular properties constituting the fingerprint. We present a scoring that takes into account dependencies between molecular properties. As before, we predict posterior probabilities of molecular properties using machine learning. Dependencies between molecular properties are modeled as a Bayesian tree network; the tree structure is estimated on the fly from the instance data. For each edge, we also estimate the expected covariance between the two random variables. For fixed marginal probabilities, we then estimate conditional probabilities using the known covariance. Now, the corrected posterior probability of each candidate can be computed, and candidates are ranked by this score. Modeling dependencies improves identification rates of CSI:FingerID by 2.85 percentage points. The new scoring Bayesian (fixed tree) is integrated into SIRIUS 4.0 (https://bio.informatik.uni-jena.de/software/sirius/).
DOI: 10.1093/nar/gks1146
发表时间: 2013-01
影响因子: 14.9
作者:
Hastings J;de Matos P;Dekker A;Ennis M;Harsha B;Kale N;Muthukrishnan V;Owen G;Turner S;Williams M;Steinbeck C
通讯作者: Steinbeck C
DOI: 10.1093/nar/gkv951
发表时间: 2016-01-04
影响因子: 14.9
作者:
Kim S;Thiessen PA;Bolton EE;Chen J;Fu G;Gindulyte A;Han L;He J;He S;Shoemaker BA;Wang J;Yu B;Zhang J;Bryant SH
通讯作者: Bryant SH
DOI: 10.1021/acs.analchem.6b00770
发表时间: 2016-08-16
影响因子: 7.4
作者:
Tsugawa H;Kind T;Nakabayashi R;Yukihira D;Tanaka W;Cajka T;Saito K;Fiehn O;Arita M
通讯作者: Arita M
DOI: 10.1021/ac400861a
发表时间: 2013-06-18
影响因子: 7.4
作者:
Ridder, Lars;van der Hooft, Justin J. J.;Vervoort, Jacques
通讯作者: Vervoort, Jacques
DOI: 10.1093/bioinformatics/btn642
发表时间: 2009-02-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Rogers, Simon;Scheltema, Richard A.;Breitling, Rainer
通讯作者: Breitling, Rainer