Expanding the understanding of biases in development of clinical-grade molecular signatures: a case study in acute respiratory viral infections.

Expanding the understanding of biases in development of clinical-grade molecular signatures: a case study in acute respiratory viral infections.
复制标题

DOI:
10.1371/journal.pone.0020662
复制
发表时间:
2011
期刊:
影响因子:
3.7
通讯作者:
Statnikov A
Statnikov A
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Lytkin NI;McVoy L;Weitkamp JH;Aliferis CF;Statnikov A

文献摘要

参考文献

被引文献

相似文献

现代个性化医疗的前景是使用分子和临床信息来更好地诊断,管理和治疗疾病,在个体患者的基础上。这些功能主要由分子特征实现,分子特征是用于从高通量测定数据预测表型和其他感兴趣的响应的计算模型。数据分析是分子特征开发的核心组成部分,如果操作不当,可能会危及整个过程。虽然探索性数据分析可以容忍次优方案,但临床级分子签名受到更严格的要求。缩小探索性与临床成功分子特征标准之间的差距需要彻底了解数据分析阶段可能存在的偏倚,并制定避免偏倚的策略。使用最近推出的数据分析协议作为案例研究,我们提供了一个深入研究的偏差的数据分析协议相关的签名的多样性,生物标志物冗余,数据预处理和验证的签名再现性。这项工作中提出的方法和结果旨在扩大对这些数据分析偏差的理解,这些数据分析偏差会影响临床稳健分子特征的发展。目前的研究提出了若干建议。首先,应尽可能提取表型的所有分子特征,以便为理解疾病发病机制提供全面和准确的依据。第二,冗余基因通常应从最终标签中去除,以促进再现性并降低制造成本。第三,数据预处理程序的设计应避免生物标志物的选择产生偏差。最后,在不同的表型和患者群体上开发和应用的分子签名应该非常谨慎地对待。
The promise of modern personalized medicine is to use molecular and clinical information to better diagnose, manage, and treat disease, on an individual patient basis. These functions are predominantly enabled by molecular signatures, which are computational models for predicting phenotypes and other responses of interest from high-throughput assay data. Data-analytics is a central component of molecular signature development and can jeopardize the entire process if conducted incorrectly. While exploratory data analysis may tolerate suboptimal protocols, clinical-grade molecular signatures are subject to vastly stricter requirements. Closing the gap between standards for exploratory versus clinically successful molecular signatures entails a thorough understanding of possible biases in the data analysis phase and developing strategies to avoid them. Using a recently introduced data-analytic protocol as a case study, we provide an in-depth examination of the poorly studied biases of the data-analytic protocols related to signature multiplicity, biomarker redundancy, data preprocessing, and validation of signature reproducibility. The methodology and results presented in this work are aimed at expanding the understanding of these data-analytic biases that affect development of clinically robust molecular signatures. Several recommendations follow from the current study. First, all molecular signatures of a phenotype should be extracted to the extent possible, in order to provide comprehensive and accurate grounds for understanding disease pathogenesis. Second, redundant genes should generally be removed from final signatures to facilitate reproducibility and decrease manufacturing costs. Third, data preprocessing procedures should be designed so as not to bias biomarker selection. Finally, molecular signatures developed and applied on different phenotypes and populations of patients should be treated with great caution.
DOI: 10.1182/blood-2006-02-002477
发表时间: 2007-03-01
期刊: BLOOD
影响因子: 20.3
作者:
Ramilo, Octavio;Allman, Windy;Chaussabel, Damien
通讯作者: Chaussabel, Damien
DOI: 10.1001/archinte.1958.00260140099015
发表时间: 1958-01-01
影响因子: --
作者:
JACKSON, GG;DOWLING, HF;BOAND, AV
通讯作者: BOAND, AV
DOI: 10.1093/bioinformatics/btg410
发表时间: 2004-02-12
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Cope, LM;Irizarry, RA;Speed, TP
通讯作者: Speed, TP
DOI: 10.1111/j.2517-6161.1995.tb02031.x
发表时间: 1995-01-01
影响因子: 5.8
作者:
BENJAMINI, Y;HOCHBERG, Y
通讯作者: HOCHBERG, Y
DOI: 10.1093/biostatistics/4.2.249
发表时间: 2003-04-01
期刊: BIOSTATISTICS
影响因子: 2.1
作者:
Irizarry, RA;Hobbs, B;Speed, TP
通讯作者: Speed, TP