Introduction to the Development and Validation of Predictive Biomarker Models from High-Throughput Data Sets

Introduction to the Development and Validation of Predictive Biomarker Models from High-Throughput Data Sets
复制标题

DOI:
10.1007/978-1-60761-580-4_15
复制
发表时间:
2010-01-01
期刊:
STATISTICAL METHODS IN MOLECULAR BIOLOGY
影响因子:
--
通讯作者:
Campagne, Fabien
Campagne, Fabien
中科院分区:
其他
文献类型:
--
作者:
Deng, Xutao;Campagne, Fabien

文献摘要

被引文献

相似文献

高通量技术可以常规地分析生物或临床样品,并产生广泛的数据集,其中每个样品与数万个测量值相关联。可以挖掘这样的数据集以发现生物标志物并开发能够从样品中测量的数据预测感兴趣的终点的统计模型。生物标志物模型开发领域结合了统计学和机器学习的方法来开发和评估预测性生物标志物模型。在本章中,我们讨论了生物标志物模型的发展所涉及的计算步骤,该模型旨在预测个体样本的信息,并回顾了通常用于实现每个步骤的方法。在一个大的基因表达数据集的生物标志物模型开发的一个实际例子。该实施例利用BDVal,作为开源项目开发的一套生物标志物模型开发程序(参见http://bdval.org/)。
High-throughput technologies can routinely assay biological or clinical samples and produce wide data sets where each sample is associated with tens of thousands of measurements. Such data sets can be mined to discover biomarkers and develop statistical models capable of predicting an endpoint of interest from data measured in the samples. The field of biomarker model development combines methods from statistics and machine learning to develop and evaluate predictive biomarker models. In this chapter, we discuss the computational steps involved in the development of biomarker models designed to predict information about individual samples and review approaches often used to implement each step. A practical example of biomarker model development in a large gene expression data set is presented. This example leverages BDVal, a suite of biomarker model development programs developed as an open-source project (see http://bdval.org/).