Research issues and strategies for genomic and proteomic biornarker discovery and validation: a statistical perspective

Research issues and strategies for genomic and proteomic biornarker discovery and validation: a statistical perspective
复制标题

DOI:
10.1517/14622416.5.6.709
复制
发表时间:
2004-09-01
期刊:
影响因子:
2.1
通讯作者:
Srivastava, S
Srivastava, S
中科院分区:
医学4区
文献类型:
--
作者:
Feng, Z;Prentice, R;Srivastava, S

文献摘要

被引文献

相似文献

从高维基因组和蛋白质组信息中开发和验证临床上有用的生物标志物构成了巨大的研究挑战。目前的瓶颈包括:在最初发现中显示出希望的生物标志物很少被发现保证随后的验证;并且生物标志物验证是昂贵且耗时的。生物标志物评价应有序进行,以提高严谨性和效率。分子谱分析方法虽然有前途,但很有可能产生有偏见的结果和过度拟合的模型。来自队列或干预试验的样本对于消除偏倚至关重要。生物标志物验证的高成本激发了一些新的研究设计特征,包括顺序过滤和DNA合并。对于数据分析,逻辑回归(特别是增强逻辑回归)具有对模型错误指定的鲁棒性,并且具有对模型过拟合的抵抗力。模型评估和交叉验证是数据分析的关键组成部分。拥有独立的测试集是研究设计的重要特征。
The development and validation of clinically useful biomarkers from high-dimensional genomic and proteomic information pose great research challenges. Present bottlenecks include: that few of the biomarkers showing promise in initial discovery were found to warrant subsequent validation; and biomarker validation is expensive and time consuming. Biomarker evaluation should proceed in an orderly fashion to enhance rigor and efficiency. A molecular profiling approach, although promising, has a high chance of yielding biased results and overfitted models. Specimens from cohorts or intervention trials are essential to eliminate biases. The high cost for biomarker validation motivates some novel study design features, including sequential filtering and DNA pooling. For data analysis, logistic regression (in particular, boosting logistic regression) has features of robustness against model misspecification, and has resistance to model overfitting. Model assessment and cross-validation are critical components of data analysis. Having an independent test set is a vital feature of study design.