Getting DNA copy numbers without control samples

Getting DNA copy numbers without control samples
复制标题

在没有对照样本的情况下获取 DNA 拷贝数

DOI:
--
复制
发表时间:
2012
影响因子:
1
通讯作者:
Á. Rubio
Á. Rubio
中科院分区:
生物学4区
文献类型:
--
作者:
M. Ortiz;Ander Aramburu;Á. Rubio

文献摘要

参考文献

被引文献

相似文献

背景在拷贝数分析中选择参考来衡量数据对于实现准确的估计是至关重要的。通常,该参考是使用研究中包括的对照样本生成的。然而,这些对照样品并不总是可用的,在这些情况下,必须创建一个人工参照。对于噪声和偏差而言,这种信号的正确产生是至关重要的。我们提出了正态搜索算法(NSA),这是一种在有或没有控制样本的情况下工作的缩放方法。它是基于这样的假设,即两个等位基因中拷贝数相同的SNPs丰富的基因组区域可能是正常的。分别为每个样本预测这些正常区域,并使用它们来计算最终参考信号。NSA可以应用于任何CN数据,而不受微阵列技术和预处理方法的影响。结果分析了5个人类数据集(HapMap样本子集、多形性胶质母细胞瘤(GBM)、卵巢、前列腺癌和肺癌实验)。结果表明,仅使用肿瘤样本,NSA就能消除拷贝数估计中的偏差,降低噪声,从而提高检测拷贝数像差(CNA)的能力。这些改进使得NSA还可以比其他最先进的方法更准确地检测反复出现的像差。结论NSA为将探头信号数据缩放到CN值提供了稳健和准确的参考,而不需要对照样本。它最大限度地减少了CNS估计中的偏差、噪声和批次效应问题。因此,与现有方法相比,NSA Scaling方法有助于更好地检测复发的CNA。自动选择引用使得执行许多GEO或ArrayExpress实验的批量分析非常有用,而无需开发解析器来查找数据中的正常样本或可能的批次。该方法可在开源R包nsa中获得,该包是aroma.cn framework.http://www.aroma-project.org/addons.的一个附加组件
BackgroundThe selection of the reference to scale the data in a copy number analysis has paramount importance to achieve accurate estimates. Usually this reference is generated using control samples included in the study. However, these control samples are not always available and in these cases, an artificial reference must be created. A proper generation of this signal is crucial in terms of both noise and bias.We propose NSA (Normality Search Algorithm), a scaling method that works with and without control samples. It is based on the assumption that genomic regions enriched in SNPs with identical copy numbers in both alleles are likely to be normal. These normal regions are predicted for each sample individually and used to calculate the final reference signal. NSA can be applied to any CN data regardless the microarray technology and preprocessing method. It also finds an optimal weighting of the samples minimizing possible batch effects.ResultsFive human datasets (a subset of HapMap samples, Glioblastoma Multiforme (GBM), Ovarian, Prostate and Lung Cancer experiments) have been analyzed. It is shown that using only tumoral samples, NSA is able to remove the bias in the copy number estimation, to reduce the noise and therefore, to increase the ability to detect copy number aberrations (CNAs). These improvements allow NSA to also detect recurrent aberrations more accurately than other state of the art methods.ConclusionsNSA provides a robust and accurate reference for scaling probe signals data to CN values without the need of control samples. It minimizes the problems of bias, noise and batch effects in the estimation of CNs. Therefore, NSA scaling approach helps to better detect recurrent CNAs than current methods. The automatic selection of references makes it useful to perform bulk analysis of many GEO or ArrayExpress experiments without the need of developing a parser to find the normal samples or possible batches within the data. The method is available in the open-source R package NSA, which is an add-on to the aroma.cn framework.http://www.aroma-project.org/addons.
DOI: 10.1093/biostatistics/kxh008
发表时间: 2004-10-01
期刊: BIOSTATISTICS
影响因子: 2.1
作者:
Olshen, AB;Venkatraman, ES;Wigler, M
通讯作者: Wigler, M
DOI: 10.1093/biostatistics/kxq043
发表时间: 2011-01-01
期刊: BIOSTATISTICS
影响因子: 2.1
作者:
Scharpf, Robert B.;Ruczinski, Ingo;Irizarry, Rafael A.
通讯作者: Irizarry, Rafael A.
DOI: 10.1101/gr.5402306
发表时间: 2006-09-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Peiffer, Daniel A.;Le, Jennie M.;Gunderson, Kevin L.
通讯作者: Gunderson, Kevin L.