MODEL SELECTION FOR HIGH-DIMENSIONAL, MULTI-SEQUENCE CHANGE-POINT PROBLEMS

MODEL SELECTION FOR HIGH-DIMENSIONAL, MULTI-SEQUENCE CHANGE-POINT PROBLEMS
复制标题

DOI:
10.5705/ss.2010.257
复制
发表时间:
2012
期刊:
影响因子:
1.4
通讯作者:
N. Zhang;D. Siegmund
N. Zhang;D. Siegmund
中科院分区:
数学3区
文献类型:
--
作者:
N. Zhang;D. Siegmund

文献摘要

被引文献

相似文献

变点模型已广泛应用于空间或时间序列数据的分割。最近在基因组学中的一些应用激发了多序列变点模型,用于跨多个比对序列的共享变化。这些应用程序经常涉及数据的变化点的数量可能很大。在以前的论文中,我们推导出贝叶斯信息准则(BIC),用于确定一个独立的正常观测序列的平均值的变化的数量时,变点的数量m被假定为保持有界的观测数量的增加。在这里,我们将该结果扩展到m可以随样本大小增加的情况,以及多个序列中的同时变点。进入新准则的随机项涉及具有负漂移的双边随机游动的积分和极大值。新的标准适用于DNA拷贝数数据的分析。
Change-point models have been widely applied for segmentation of spatial or time-series data. Some recent applications in genomics motivate multi-sequence change-point models for shared changes across multiple aligned sequences. These applications frequently involve data where the number of change-points can be large. In a previous paper we derived a Bayes Information Criterion (BIC) for determining the number of changes in the mean of a sequence of independent normal observations when the number of change-points m is assumed to remain bounded as the number of observations increases. Here we extend that result to the case where m can increase with the sample size and to simultaneous change-points in multiple sequences. Stochastic terms that enter into the new criteria involve integrals and maxima of two-sided random walks with negative drift. The new criteria are applied to the analysis of DNA copy number data.