A fast Bayesian change point analysis for the segmentation of microarray data

A fast Bayesian change point analysis for the segmentation of microarray data
复制标题

DOI:
10.1093/bioinformatics/btn404
复制
发表时间:
2008-10-01
期刊:
影响因子:
5.8
通讯作者:
Emerson, John W.
Emerson, John W.
中科院分区:
生物学3区
文献类型:
--
作者:
Erdman, Chandra;Emerson, John W.

文献摘要

被引文献

相似文献

动机:检测基因改变区域的能力在癌症研究中非常重要。这些改变可以采取大的染色体获得和丢失以及较小的扩增和缺失的形式。这些区域的检测使研究人员能够识别参与癌症进展的基因,并充分了解癌症和非癌症组织之间的差异。由巴里和Hartigan提出的贝叶斯方法非常适合于分析这种变点问题。在上一篇文章中,我们介绍了R包bcp(贝叶斯变点),这是巴里和哈迪根方法的MCMC实现。在模拟研究和真实的数据例子中,bcp被证明可以准确地检测变化点和估计段均值。bcp的早期版本(2.0之前)在速度上是O(n(2)),在内存上是O(n)(其中n是观察的数量),并且对于长度为10000的序列,运行时间类似于45 min。随着新的微阵列的高分辨率,在O(n2)算法的计算数量是令人望而却步的时间intensive.Results:我们提出了一个新的实现的贝叶斯变点方法,是O(n)的速度和内存; BCP 2.1运行在一个单一的处理器上类似于45秒的长度为10 000的序列-一个巨大的速度增益。使用并行计算可以进一步提高速度,通过NetWorkSpaces在bcp中提供支持。在来自文献的模拟和真实的微阵列数据中,bcp被证明可以快速准确地检测不同宽度和大小的畸变。
Motivation: The ability to detect regions of genetic alteration is of great importance in cancer research. These alterations can take the form of large chromosomal gains and losses as well as smaller amplications and deletions. The detection of such regions allows researchers to identify genes involved in cancer progression, and to fully understand differences between cancer and non-cancer tissue. The Bayesian method proposed by Barry and Hartigan is well suited for the analysis of such change point problems. In our previous article we introduced the R package bcp (Bayesian change point), an MCMC implementation of Barry and Hartigan's method. In a simulation study and real data examples, bcp is shown to both accurately detect change points and estimate segment means. Earlier versions of bcp (prior to 2.0) are O(n(2)) in speed and O(n) in memory (where n is the number of observations), and run in similar to 45 min for a sequence of length 10 000. With the high resolution of newer microarrays, the number of computations in the O(n2) algorithm is prohibitively time-intensive.Results: We present a new implementation of the Bayesian change point method that is O(n) in both speed and memory; bcp 2.1 runs in similar to 45s on a single processor with a sequence of length 10 000-a tremendous speed gain. Further speed improvements are possible using parallel computing, supported in bcp via NetWorkSpaces. In simulated and real microarray data from the literature, bcp is shown to quickly and accurately detect aberrations of varying width and magnitude.