Breast cancer prognosis by combinatorial analysis of gene expression data.

Breast cancer prognosis by combinatorial analysis of gene expression data.
复制标题

DOI:
10.1186/bcr1512
复制
发表时间:
2006
期刊:
Breast cancer research : BCR
影响因子:
--
通讯作者:
Hammer PL
Hammer PL
中科院分区:
其他
文献类型:
--
作者:
Alexe G;Alexe S;Axelrod DE;Bonates TO;Lozina II;Reiss M;Hammer PL

文献摘要

被引文献

相似文献

van 't Veer及其同事最近的乳腺癌数据集说明了将数据分析工具应用于微阵列数据的诊断和预后的潜力。我们使用数据逻辑分析(LAD)的新技术重新检查该数据集,其双重目标是发现具有良好或不良结果的案例的模式特征,并将其用于准确和合理的预测;并得到关于基因作用的新信息,特殊类型病例的存在,以及其他因素。数据分析使用基于组合和优化的LAD方法,最近显示在心脏病学、癌症蛋白质组学、血液学、肺脏学和其他学科中提供高度准确的诊断和预后系统。LAD鉴定出25000个基因中的17个子集,能够完全区分预后较差和良好的患者。生成了一个广泛的“模式”或“组合生物标志物”列表(即基因的组合及其表达水平的限制),并使用40种模式创建了一个预后系统,在训练集和测试集上分别显示出100%和92.9%的加权准确率。与其他方法相比,该预测系统使用的基因更少,并且与其他研究报告的准确性相似或更高。在LAD鉴定的17个基因中,3个(分别为5个)被证明在决定不良预后(分别为良好预后)中起重要作用。发现了两类新的患者(由相似的覆盖模式、基因表达范围和临床特征描述)。研究结果表明,van 't Veer的训练集和测试集具有不同的特征。本研究表明,LAD利用基因组数据为乳腺癌提供了一个准确且充分解释的预后系统(即,该系统除了预测预后好坏外,还为每位患者提供了对预后原因的个性化解释)。此外,LAD模型对个体和组合生物标志物的作用提供了有价值的见解,允许发现新的患者类别,并产生了大量的生物医学研究假设库。
The potential of applying data analysis tools to microarray data for diagnosis and prognosis is illustrated on the recent breast cancer dataset of van 't Veer and coworkers. We re-examine that dataset using the novel technique of logical analysis of data (LAD), with the double objective of discovering patterns characteristic for cases with good or poor outcome, using them for accurate and justifiable predictions; and deriving novel information about the role of genes, the existence of special classes of cases, and other factors. Data were analyzed using the combinatorics and optimization-based method of LAD, recently shown to provide highly accurate diagnostic and prognostic systems in cardiology, cancer proteomics, hematology, pulmonology, and other disciplines. LAD identified a subset of 17 of the 25,000 genes, capable of fully distinguishing between patients with poor, respectively good prognoses. An extensive list of 'patterns' or 'combinatorial biomarkers' (that is, combinations of genes and limitations on their expression levels) was generated, and 40 patterns were used to create a prognostic system, shown to have 100% and 92.9% weighted accuracy on the training and test sets, respectively. The prognostic system uses fewer genes than other methods, and has similar or better accuracy than those reported in other studies. Out of the 17 genes identified by LAD, three (respectively, five) were shown to play a significant role in determining poor (respectively, good) prognosis. Two new classes of patients (described by similar sets of covering patterns, gene expression ranges, and clinical features) were discovered. As a by-product of the study, it is shown that the training and the test sets of van 't Veer have differing characteristics. The study shows that LAD provides an accurate and fully explanatory prognostic system for breast cancer using genomic data (that is, a system that, in addition to predicting good or poor prognosis, provides an individualized explanation of the reasons for that prognosis for each patient). Moreover, the LAD model provides valuable insights into the roles of individual and combinatorial biomarkers, allows the discovery of new classes of patients, and generates a vast library of biomedical research hypotheses.