Identifying multiple changepoints in heterogeneous binary data with an application to molecular genetics

Identifying multiple changepoints in heterogeneous binary data with an application to molecular genetics
复制标题

DOI:
10.1093/biostatistics/kxh005
复制
发表时间:
2004-10-01
期刊:
影响因子:
2.1
通讯作者:
Taylor, PR
Taylor, PR
中科院分区:
数学2区
文献类型:
--
作者:
Albert, PS;Hunsberger, SA;Taylor, PR

文献摘要

被引文献

相似文献

突变点的识别是分子遗传学中的一个重要问题。我们的激励例子来自癌症遗传学,其兴趣集中在确定染色体中肿瘤抑制基因可能性增加的区域。杂合性缺失(LOH)是等位基因缺失的一种二元测量方法,其中LOH频率沿染色体的突变可以识别表明含有肿瘤抑制基因的区域的边界。我们的兴趣是测试多个变化点的存在,以确定LOH频率增加的区域。一个复杂的因素是患者之间LOH频率的巨大异质性,其中一些患者的LOH频率非常高,而另一些患者的LOH频率较低。我们开发了一个程序来识别异构二进制数据中的多个变化点。我们提出了近似和完全最大似然方法,并将这两种方法与忽略二元数据异质性的朴素方法进行比较。该方法用于估计食管癌患者13号染色体上LOH频率的模式,并分离出13号染色体上可能含有肿瘤抑制基因的LOH频率过高的区域。通过模拟,我们表明我们的方法工作得很好,并且对于偏离一些关键的建模假设是稳健的。
Identifying changepoints is an important problem in molecular genetics. Our motivating example is from cancer genetics where interest focuses on identifying areas of a chromosome with an increased likelihood of a tumor suppressor gene. Loss of heterozygosity (LOH) is a binary measure of allelic loss in which abrupt changes in LOH frequency along the chromosome may identify boundaries indicative of a region containing a tumor suppressor gene. Our interest was on testing for the presence of multiple changepoints in order to identify regions of increased LOH frequency. A complicating factor is the substantial heterogeneity in LOH frequency across patients, where some patients have a very high LOH frequency while others have a low frequency. We develop a procedure for identifying multiple changepoints in heterogeneous binary data. We propose both approximate and full maximum-likelihood approaches and compare these two approaches with a naive approach in which we ignore the heterogeneity in the binary data. The methodology is used to estimate the pattern in LOH frequency on chromosome 13 in esophageal cancer patients and to isolate an area of inflated LOH frequency on chromosome 13 which may contain a tumor suppressor gene. Using simulations, we show that our approach works well and that it is robust to departures from some key modeling assumptions.