Sparse representation and Bayesian detection of genome copy number alterations from microarray data

Sparse representation and Bayesian detection of genome copy number alterations from microarray data
复制标题

DOI:
10.1093/bioinformatics/btm601
复制
发表时间:
2008-02-01
期刊:
影响因子:
5.8
通讯作者:
Asgharzadeh, Shahab
Asgharzadeh, Shahab
中科院分区:
生物学3区
文献类型:
--
作者:
Pique-Regi, Roger;Monso-Varona, Jordi;Asgharzadeh, Shahab

文献摘要

被引文献

相似文献

动机:癌症中的基因组不稳定性导致与肿瘤的发展和行为相关的异常基因组拷贝数改变(CNA)。微阵列技术的进步已经允许在检测基因组中的DNA拷贝数变化(扩增或缺失)中具有更高的分辨率。然而,来自阵列探针的测量信号和伴随的噪声的数量的增加在定义CNA的断点的准确和快速识别方面提出了挑战。本文提出了一种新的检测技术,利用分段常数(PWC)向量表示基因组拷贝数和稀疏贝叶斯学习(SBL)检测CNA breakpoints.Methods:首先,一个紧凑的线性代数表示的基因组拷贝数是从归一化探针强度。第二,应用并优化SBL以推断拷贝数变化发生的位置。第三,一个向后消除(BE)的过程是用来排名推断的断点;和一个截止点可以有效地调整在这个过程中,以控制错误发现率(FDR)。结果:我们的算法的性能进行评估,使用模拟和真实的基因组数据集,并与其他现有的技术相比。我们的方法实现了最高的准确度和最低的FDR,同时将计算速度提高了几个数量级。所提出的算法已被开发成一个独立的软件应用程序(GADA,基因组变异检测算法)。
Motivation: Genomic instability in cancer leads to abnormal genome copy number alterations (CNA) that are associated with the development and behavior of tumors. Advances in microarray technology have allowed for greater resolution in detection of DNA copy number changes (amplifications or deletions) across the genome. However, the increase in number of measured signals and accompanying noise from the array probes present a challenge in accurate and fast identification of breakpoints that define CNA. This article proposes a novel detection technique that exploits the use of piece wise constant (PWC) vectors to represent genome copy number and sparse Bayesian learning (SBL) to detect CNA breakpoints.Methods: First, a compact linear algebra representation for the genome copy number is developed from normalized probe intensities. Second, SBL is applied and optimized to infer locations where copy number changes occur. Third, a backward elimination (BE) procedure is used to rank the inferred breakpoints; and a cut-off point can be efficiently adjusted in this procedure to control for the false discovery rate (FDR).Results: The performance of our algorithm is evaluated using simulated and real genome datasets and compared to other existing techniques. Our approach achieves the highest accuracy and lowest FDR while improving computational speed by several orders of magnitude. The proposed algorithm has been developed into a free standing software application (GADA, Genome Alteration Detection Algorithm).