Predicting the prognosis of breast cancer by integrating clinical and microarray data with Bayesian networks

Predicting the prognosis of breast cancer by integrating clinical and microarray data with Bayesian networks
复制标题

DOI:
10.1093/bioinformatics/btl230
复制
发表时间:
2006-07-01
期刊:
影响因子:
5.8
通讯作者:
De Moor, Bart
De Moor, Bart
中科院分区:
生物学3区
文献类型:
--
作者:
Gevaert, Olivier;De Smet, Frank;De Moor, Bart

文献摘要

被引文献

相似文献

动机:临床数据,如病史,实验室分析,超声参数,这是日常的临床决策支持的基础,往往是未充分利用,以指导临床管理的癌症微阵列数据的存在。我们提出了一种基于贝叶斯网络的策略,平等对待临床和微阵列数据。这种概率模型的主要优点是,它允许以多种方式集成这些数据源,并且允许调查和理解模型结构和参数。此外,使用马尔可夫毯的概念,我们可以识别所有的变量,这些变量屏蔽了类变量的影响,其余的网络。因此,贝叶斯网络自动执行特征选择,通过识别(中)的依赖关系与类variable.Results:我们评估了三种方法整合临床和微阵列数据:决策集成,部分集成和完全集成,并使用它们来分类公开的数据对乳腺癌患者的预后不良和良好的一组。部分积分方法是最有前途的,并具有独立的测试集面积0.845的ROC曲线下。在选择一个工作点后,分类性能优于常用的指数。
Motivation: Clinical data, such as patient history, laboratory analysis, ultrasound parameters-which are the basis of day-to-day clinical decision support-are often underused to guide the clinical management of cancer in the presence of microarray data. We propose a strategy based on Bayesian networks to treat clinical and microarray data on an equal footing. The main advantage of this probabilistic model is that it allows to integrate these data sources in several ways and that it allows to investigate and understand the model structure and parameters. Furthermore using the concept of a Markov Blanket we can identify all the variables that shield off the class variable from the influence of the remaining network. Therefore Bayesian networks automatically perform feature selection by identifying the ( in) dependency relationships with the class variable.Results: We evaluated three methods for integrating clinical and microarray data: decision integration, partial integration and full integration and used them to classify publicly available data on breast cancer patients into a poor and a good prognosis group. The partial integration method is most promising and has an independent test set area under the ROC curve of 0.845. After choosing an operating point the classification performance is better than frequently used indices.