Analysis of array CGH data for cancer studies using fused quantile regression

Analysis of array CGH data for cancer studies using fused quantile regression
复制标题

DOI:
10.1093/bioinformatics/btm364
复制
发表时间:
2007-09-15
期刊:
影响因子:
5.8
通讯作者:
Zhu, Ji
Zhu, Ji
中科院分区:
生物学3区
文献类型:
--
作者:
Li, Youjuan;Zhu, Ji

文献摘要

被引文献

相似文献

动机:DNA 拷贝数变化的识别提供了可能增进我们对癌症发生和进展的理解的见解。基于阵列的比较基因组杂交(阵列-CGH)已成为一种允许高通量全基因组扫描染色体畸变的技术。已经提出了许多统计方法来分析阵列 CGH 数据。在本文中,我们考虑基于三个动机的融合分位数回归模型:(1)分位数回归可以比标准均值回归方法提供更全面的拷贝数比率概况; (2)为了简单起见,大多数可用方法假设相邻克隆之间的间距均匀,而合并克隆的物理位置信息可能会有所帮助,并且(3)大多数当前方法都有一组必须仔细调整的调整参数,这给实现带来了复杂性。结果:我们在融合正则化分位数回归框架中制定了增益和损失区域的检测,合并了克隆的物理位置。我们推导了一种有效的算法,可以计算最终优化问题的整个解决方案路径,并且我们提出了对拟合模型复杂性的简单估计,从而可以方便地选择调整参数。三个已发布的阵列 CGH 数据集用于演示我们的方法。
Motivation: The identification of DNA copy number changes provides insights that may advance our understanding of initiation and progression of cancer. Array-based comparative genomic hybridization (array-CGH) has emerged as a technique allowing high-throughput genome-wide scanning for chromosomal aberrations. A number of statistical methods have been proposed for the analysis of array-CGH data. In this article, we consider a fused quantile regression model based on three motivations: (1) quantile regression may provide a more comprehensive picture for the ratio profile of copy numbers than the standard mean regression approach; (2) for simplicity, most available methods assume uniform spacing between neighboring clones, while incorporating the information of physical locations of clones may be helpful and (3) most current methods have a set of tuning parameters that must be carefully tuned, which introduces complexity to the implementation.Results: We formulate the detection of regions of gains and losses in a fused regularized quantile regression framework, incorporating physical locations of clones. We derive an efficient algorithm that computes the entire solution path for the resulting optimization problem, and we propose a simple estimate for the complexity of the fitted model, which leads to convenient selection of the tuning parameter. Three published array-CGH datasets are used to demonstrate our approach.