Kernel approaches for differential expression analysis of mass spectrometry-based metabolomics data.

Kernel approaches for differential expression analysis of mass spectrometry-based metabolomics data.
复制标题

DOI:
10.1186/s12859-015-0506-3
复制
发表时间:
2015-03-11
期刊:
影响因子:
3
通讯作者:
Ghosh D
Ghosh D
中科院分区:
生物学4区
文献类型:
--
作者:
Zhan X;Patterson AD;Ghosh D

文献摘要

参考文献

相似文献

从代谢组学实验产生的数据不同于其他类型的“组学”数据。例如,基于质谱(MS)的代谢组学数据中的一个常见现象是数据矩阵经常包含缺失值,这使一些定量分析变得复杂。解决这个问题的一个方法是把他们当作缺席。因此,代谢组学数据中有两种类型的信息:代谢物的存在/不存在以及代谢物丰度水平的定量值(如果存在)。结合这两层信息对传统统计方法在差异表达分析中的应用提出了挑战。在这篇文章中,我们提出了一种新的基于核的评分测试代谢组学差异表达分析。为了同时捕获代谢组学数据中的连续模式和离散模式,设计了两种新的核函数。一种是基于距离的核,另一种是分层核。虽然我们最初描述的程序在单代谢物分析的情况下,我们扩展的方法来处理代谢物集以及。基于模拟数据和肝癌代谢组学研究的真实的数据的评估表明,我们的核方法具有更好的性能比一些现有的替代品。在R统计计算环境中所提出的核方法的实现可在http://works.bepress.com/debashis_ghosh/60/获得。本文的在线版本(doi:10.1186/s12859-015-0506-3)包含补充材料,可供授权用户使用。
Data generated from metabolomics experiments are different from other types of “-omics” data. For example, a common phenomenon in mass spectrometry (MS)-based metabolomics data is that the data matrix frequently contains missing values, which complicates some quantitative analyses. One way to tackle this problem is to treat them as absent. Hence there are two types of information that are available in metabolomics data: presence/absence of a metabolite and a quantitative value of the abundance level of a metabolite if it is present. Combining these two layers of information poses challenges to the application of traditional statistical approaches in differential expression analysis. In this article, we propose a novel kernel-based score test for the metabolomics differential expression analysis. In order to simultaneously capture both the continuous pattern and discrete pattern in metabolomics data, two new kinds of kernels are designed. One is the distance-based kernel and the other is the stratified kernel. While we initially describe the procedures in the case of single-metabolite analysis, we extend the methods to handle metabolite sets as well. Evaluation based on both simulated data and real data from a liver cancer metabolomics study indicates that our kernel method has a better performance than some existing alternatives. An implementation of the proposed kernel method in the R statistical computing environment is available at http://works.bepress.com/debashis_ghosh/60/. The online version of this article (doi:10.1186/s12859-015-0506-3) contains supplementary material, which is available to authorized users.
DOI: 10.1186/1471-2105-7-530
发表时间: 2006-12-13
期刊: BMC bioinformatics
影响因子: 3
作者:
Baran R;Kochi H;Saito N;Suematsu M;Soga T;Nishioka T;Robert M;Tomita M
通讯作者: Tomita M
DOI: 10.2307/2335690
发表时间: 1977-01-01
期刊: BIOMETRIKA
影响因子: 2.7
作者:
DAVIES, RB
通讯作者: DAVIES, RB
DOI: 10.1074/jbc.m601876200
发表时间: 2006-06-16
影响因子: 4.8
作者:
Soga, Tomoyoshi;Baran, Richard;Tomita, Masaru
通讯作者: Tomita, Masaru
DOI: 10.1111/j.1541-0420.2007.00799.x
发表时间: 2007-12-01
期刊: BIOMETRICS
影响因子: 1.9
作者:
Liu, Dawei;Lin, Xihong;Ghosh, Debashis
通讯作者: Ghosh, Debashis
DOI: 10.1093/biomet/74.1.33
发表时间: 1987-03-01
期刊: BIOMETRIKA
影响因子: 2.7
作者:
DAVIES, RB
通讯作者: DAVIES, RB