Robust and efficient identification of biomarkers by classifying features on graphs

Robust and efficient identification of biomarkers by classifying features on graphs
复制标题

DOI:
10.1093/bioinformatics/btn383
复制
发表时间:
2008-09-15
期刊:
影响因子:
5.8
通讯作者:
Kuang, Rui
Kuang, Rui
中科院分区:
生物学3区
文献类型:
--
作者:
Hwang, TaeHyun;Sicotte, Hugues;Kuang, Rui

文献摘要

被引文献

相似文献

动机:从大规模基因表达或单核苷酸多态(SNP)数据中发现生物标记物的一个中心问题是考虑到所有特征之间的相关性的计算挑战。忽略这种相关性的方法通常在独立数据集中识别不可重现的生物标记物。提出了一种新的基于图的半监督特征分类算法,通过对二部图的学习来识别疾病标志物。该算法通过网络传播将二部图中的特征节点直接分类为正、负或中性,通过探索图中的双簇结构来捕捉样本与特征(临床变量和遗传变量)之间的依赖关系。我们的算法有两个特点:(1)我们的算法能够找到一个全局最优的标记来捕捉所有特征之间的相关性,从而在独立的微阵列或其他高吞吐量数据集上产生高度可重复性的结果;(2)我们的算法能够处理数十万个特征,因此,对于从高通量基因表达和SNP数据中识别生物标记物特别有用。此外,虽然我们的算法是为分类特征而设计的,但它也可以同时对用于疾病预后/诊断的测试样本进行分类。结果:我们应用网络传播算法研究了三个大规模乳腺癌数据集。与支持向量机和其他基线方法相比,我们的算法获得了具有竞争力的分类性能,并识别了几个与疾病具有临床或生物学相关性的标记。更重要的是,我们的算法还从独立的数据集中识别出了重复性很高的标记基因,并丰富了功能。
Motivation: A central problem in biomarker discovery from large-scale gene expression or single nucleotide polymorphism (SNP) data is the computational challenge of taking into account the dependence among all the features. Methods that ignore the dependence usually identify non-reproducible biomarkers across independent datasets. We introduce a new graph-based semi-supervised feature classification algorithm to identify discriminative disease markers by learning on bipartite graphs. Our algorithm directly classifies the feature nodes in a bipartite graph as positive, negative or neutral with network propagation to capture the dependence among both samples and features (clinical and genetic variables) by exploring bi-cluster structures in a graph. Two features of our algorithm are: (1) our algorithm can find a global optimal labeling to capture the dependence among all the features and thus, generates highly reproducible results across independent microarray or other high-thoughput datasets, (2) our algorithm is capable of handling hundreds of thousands of features and thus, is particularly useful for biomarker identification from high-throughput gene expression and SNP data. In addition, although designed for classifying features, our algorithm can also simultaneously classify test samples for disease prognosis/diagnosis.Results: We applied the network propagation algorithm to study three large-scale breast cancer datasets. Our algorithm achieved competitive classification performance compared with SVMs and other baseline methods, and identified several markers with clinical or biological relevance with the disease. More importantly, our algorithm also identified highly reproducible marker genes and enriched functions from the independent datasets.