Detecting simultaneous changepoints in multiple sequences

Detecting simultaneous changepoints in multiple sequences
复制标题

DOI:
10.1093/biomet/asq025
复制
发表时间:
2010-09-01
期刊:
影响因子:
2.7
通讯作者:
Li, Jun Z.
Li, Jun Z.
中科院分区:
数学2区
文献类型:
--
作者:
Zhang, Nancy R.;Siegmund, David O.;Li, Jun Z.

文献摘要

被引文献

相似文献

我们讨论了在多个一维噪声序列中的同一位置发生的局部信号的检测,特别注意可能只发生在一小部分序列中的相对较弱的信号。我们提出了简单的扫描和分割算法的基础上的卡方统计量为每个单独的样本,这是相当于广义似然比的模型,其中每个样本中的错误是独立的。简单的几何统计,使我们能够得到准确的解析近似的显着性水平,这样的扫描。该模型的制定是出于在多个样本中检测复发性DNA拷贝数变异的生物学问题。我们使用重复和亲子比较表明,跨样本汇集数据可以更准确地检测拷贝数变异。我们还将多样本分割算法应用于分析一组包含复杂嵌套和重叠拷贝数畸变的肿瘤样本,我们的方法给出了稀疏和直观的交叉样本摘要。
We discuss the detection of local signals that occur at the same location in multiple one-dimensional noisy sequences, with particular attention to relatively weak signals that may occur in only a fraction of the sequences. We propose simple scan and segmentation algorithms based on the sum of the chi-squared statistics for each individual sample, which is equivalent to the generalized likelihood ratio for a model where the errors in each sample are independent. The simple geometry of the statistic allows us to derive accurate analytic approximations to the significance level of such scans. The formulation of the model is motivated by the biological problem of detecting recurrent DNA copy number variants in multiple samples. We show using replicates and parent-child comparisons that pooling data across samples results in more accurate detection of copy number variants. We also apply the multisample segmentation algorithm to the analysis of a cohort of tumour samples containing complex nested and overlapping copy number aberrations, for which our method gives a sparse and intuitive cross-sample summary.