Power Analysis and Sample Size Determination in Metabolic Phenotyping

Power Analysis and Sample Size Determination in Metabolic Phenotyping
复制标题

DOI:
10.1021/acs.analchem.6b00188
复制
发表时间:
2016-05-17
影响因子:
7.4
通讯作者:
Ebbels, Timothy M. D.
Ebbels, Timothy M. D.
中科院分区:
化学1区
文献类型:
--
作者:
Blaise, Benjamin J.;Correia, Goncalo;Ebbels, Timothy M. D.

文献摘要

被引文献

相似文献

统计功效和样本量的估计是实验设计的一个关键方面。然而,在代谢表型分析中,目前还没有可接受的方法来完成这些任务,这在很大程度上是由于预期效果的未知性质。在这种无假设的科学中,既不知道重要分析物的数量或类别,也不知道效应大小。我们介绍了一种新的方法,基于多变量模拟,有效地处理高度相关的结构和高维的代谢表型数据。首先,一个大的数据集模拟的基础上的一个试点研究调查一个给定的生物医学问题的特点。然后添加给定大小的效应,对应于离散(分类)或连续(回归)结果。通过从模拟数据中随机选择各种大小的数据集来建模不同的样本大小。我们研究了不同的方法进行效果检测,包括单变量和多变量技术。我们的框架使我们能够调查样本量,功率和效应量之间的复杂关系的真实的多元数据集。例如,我们证明了一个示例试点数据集,某些特征在20个样本的样本量下达到0.8的功效,或者在0.2和200个样本的效应量下达到0.8的交叉验证预测性Q(Y)(2)。我们对来自人类和模式生物C的核磁共振和液相色谱-质谱数据的方法进行了验证。优美的
Estimation of statistical power and sample size is a key aspect of experimental design. However, in metabolic phenotyping, there is currently no accepted approach for these tasks, in large part due to the unknown nature of the expected effect. In such hypothesis free science, neither the number or class of important analytes nor the effect size are known a priori. We introduce a new approach, based on multivariate simulation, which deals effectively with the highly correlated structure and high-dimensionality of metabolic phenotyping data. First, a large data set is simulated based on the characteristics of a pilot study investigating a given biomedical issue. An effect of a given size, corresponding either to a discrete (classification) or continuous (regression) outcome is then added. Different sample sizes are modeled by randomly selecting data sets of various sizes from the simulated data. We investigate different methods for effect detection, including univariate and multivariate techniques. Our framework allows us to investigate the complex relationship between sample size, power, and effect size for real multivariate data sets. For instance, we demonstrate for an example pilot data set that certain features achieve a power of 0.8 for a sample size of 20 samples or that a cross-validated predictivity Q(Y)(2) of 0.8 is reached with an effect size of 0.2 and 200 samples. We exemplify the approach for both nuclear magnetic resonance and liquid chromatography-mass spectrometry data from humans and the model organism C. elegans.