Flexible and accurate detection of genomic copy-number changes from aCGH.

Flexible and accurate detection of genomic copy-number changes from aCGH.
复制标题

ACGH的基因组拷贝数变化的灵活,准确检测。

DOI:
10.1371/journal.pcbi.0030122
复制
发表时间:
2007-06
影响因子:
4.3
通讯作者:
Diaz-Uriarte, Ramon
Diaz-Uriarte, Ramon
中科院分区:
生物学2区
文献类型:
--
作者:
Rueda, Oscar M.;Diaz-Uriarte, Ramon

文献摘要

参考文献

被引文献

相似文献

基因组 DNA 拷贝数改变 (CNA) 与包括癌症在内的复杂疾病相关:CNA 确实与肿瘤分级、转移和患者生存相关。从基于阵列的比较基因组杂交 (aCGH) 数据中发现的 CNA 有助于识别疾病相关基因和潜在的治疗靶点。为了立即在临床和基础研究场景中发挥作用,aCGH 数据分析需要准确的方法,这些方法不会强加不切实际的生物学假设,并为关键问题“该基因/区域具有 CNA 的概率是多少?”提供直接答案。然而,当前的方法无法满足这些要求。在这里,我们引入可逆跳转aCGH(RJaCGH),这是一种从aCGH中识别CNA的新方法;我们使用通过可逆跳跃马尔可夫链蒙特卡罗拟合的非齐次隐马尔可夫模型;我们通过贝叶斯模型平均纳入模型不确定性。 RJaCGH 提供了基因/区域具有 CNA 的概率估计,同时结合了探针间距离以及在染色体或全基因组基础上分析数据的能力。 RJaCGH 优于替代方法,并且对于噪声数据和高度可变的探针间距离(aCGH 数据中常见的特征),性能差异甚至更大。此外,我们的概率方法使我们能够识别样本中 CNA 的最小公共区域,并且可以扩展以合并表达数据。总之,我们提供了一个严格的统计框架,用于使用 CNA 定位基因和染色体区域,并有可能应用于癌症和其他复杂的人类疾病。由于细胞分裂过程中出现问题,染色体中基因的拷贝数可能会增加或减少。这些拷贝数改变(CNA)在复杂多基因疾病的出现中发挥着至关重要的作用。例如,在癌症中,癌基因的扩增可以驱动肿瘤激活,而 CNA 与转移发展和患者生存相关。最近,基于阵列的比较基因组杂交 (aCGH) 的广泛使用推动了对 CNA 与疾病之间关系的研究,这种技术比以前的实验方法具有更精细的分辨率。从这些数据中检测 CNA 取决于分析方法,这些方法不会强加生物学上不切实际的假设,并为基础研究问题提供直接答案。我们开发了一种统计方法,使用贝叶斯方法,从 aCGH 数据返回 CNA 概率的估计,这是对关键生物学问题的最直接和最有价值的答案:“这个基因/区域具有改变的拷贝数的概率是多少?”因此,该方法的输出可以立即用于从临床到基础研究场景的不同设置,并且适用于各种 aCGH 技术。
Genomic DNA copy-number alterations (CNAs) are associated with complex diseases, including cancer: CNAs are indeed related to tumoral grade, metastasis, and patient survival. CNAs discovered from array-based comparative genomic hybridization (aCGH) data have been instrumental in identifying disease-related genes and potential therapeutic targets. To be immediately useful in both clinical and basic research scenarios, aCGH data analysis requires accurate methods that do not impose unrealistic biological assumptions and that provide direct answers to the key question, “What is the probability that this gene/region has CNAs?” Current approaches fail, however, to meet these requirements. Here, we introduce reversible jump aCGH (RJaCGH), a new method for identifying CNAs from aCGH; we use a nonhomogeneous hidden Markov model fitted via reversible jump Markov chain Monte Carlo; and we incorporate model uncertainty through Bayesian model averaging. RJaCGH provides an estimate of the probability that a gene/region has CNAs while incorporating interprobe distance and the capability to analyze data on a chromosome or genome-wide basis. RJaCGH outperforms alternative methods, and the performance difference is even larger with noisy data and highly variable interprobe distance, both commonly found features in aCGH data. Furthermore, our probabilistic method allows us to identify minimal common regions of CNAs among samples and can be extended to incorporate expression data. In summary, we provide a rigorous statistical framework for locating genes and chromosomal regions with CNAs with potential applications to cancer and other complex human diseases. As a consequence of problems during cell division, the number of copies of a gene in a chromosome can either increase or decrease. These copy-number alterations (CNAs) can play a crucial role in the emergence of complex multigenic diseases. For example, in cancer, amplification of oncogenes can drive tumor activation, and CNAs are associated with metastasis development and patient survival. Studies on the relationship between CNAs and disease have been recently fueled by the widespread use of array-based comparative genomic hybridization (aCGH), a technique with much finer resolution than previous experimental approaches. Detection of CNAs from these data depends on methods of analysis that do not impose biologically unrealistic assumptions and that provide direct answers to fundamental research questions. We have developed a statistical method, using a Bayesian approach, that returns estimates of the probabilities of CNAs from aCGH data, the most direct and valuable answer to the key biological question: “What is the probability that this gene/region has an altered copy number?” The output of the method can therefore be immediately used in different settings from clinical to basic research scenarios, and is applicable over a wide variety of aCGH technologies.
DOI: 10.1093/bioinformatics/btl089
发表时间: 2006-05-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Marioni, JC;Thorne, NP;Tavaré, S
通讯作者: Tavaré, S
DOI: 10.1002/gcc.20382
发表时间: 2007-01-01
影响因子: 3.7
作者:
Habermann, Jens K.;Paulsen, Ulrike;Ried, Thomas
通讯作者: Ried, Thomas
DOI: 10.1093/bioinformatics/bti646
发表时间: 2005-10-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Huang, T;Wu, BL;Zhao, HY
通讯作者: Zhao, HY
DOI: 10.2307/1390675
发表时间: 1998-12-01
影响因子: 2.4
作者:
Brooks, SP;Gelman, A
通讯作者: Gelman, A
DOI: 10.1093/bioinformatics/btl035
发表时间: 2006-04-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Broët, P;Richardson, S
通讯作者: Richardson, S