Software comparison for evaluating genomic copy number variation for Affymetrix 6.0 SNP array platform.

Software comparison for evaluating genomic copy number variation for Affymetrix 6.0 SNP array platform.
复制标题

DOI:
10.1186/1471-2105-12-220
复制
发表时间:
2011-05-31
期刊:
影响因子:
3
通讯作者:
de Andrade M
de Andrade M
中科院分区:
生物学4区
文献类型:
--
作者:
Eckel-Passow JE;Atkinson EJ;Maharjan S;Kardia SL;de Andrade M

文献摘要

参考文献

被引文献

相似文献

拷贝数数据通常是从使用各种软件的全基因组关联研究芯片中提取出来的。我们对Affymetrix SNP芯片设计的四个免费软件包进行了实证评估和比较,以估计拷贝数:Affymetrix Power Tools (APT), Aroma。Affymetrix, PennCNV和CRLMM。我们的评估使用了1418份热那亚样本,这些样本在Affymetrix全基因组人类SNP阵列6.0上进行了基因分型。我们比较了基因座水平拷贝数数据的偏倚和方差,拷贝数增加/删除区域之间的一致性以及删除片段之间的假阳性率。APT的中位基因拷贝数最接近2,而PennCNV和Aroma的中位基因拷贝数最接近2。Affymetrix与中位拷贝数相关的变异性最小。在这些评估中,只有PennCNV提供了拷贝数特定的质量控制指标,并确定了136个不良的CNV样本。利用PennCNV和CRLMM/VanillaIce提供的隐马尔可夫模型检测拷贝数变异区域(CNV)。PennCNV比CRLMM/ vanilla检测到更多的cnv;每个样本检测到的CNVs中位数分别为39和30。PennCNV检测到CRLMM/VanillaIce检测到的大部分区域以及额外的CNV区域。PennCNV与CRLMM/VanillaIce的重复一致性中位数为47.9%,缺失一致性中位数为51.5%。与缺失相关的估计假阳性率在PennCNV和CRLMM/VanillaIce中相似。如果目标是对基因座级别的拷贝数数据进行统计测试,我们的经验结果表明,PennCNV或Aroma。Affymetrix是最优的。如果目标是对汇总的分段数据进行统计测试,那么PennCNV将优于CRLMM/ vanilla。具体来说,PennCNV允许分析人员在单个软件包中估计位点级别的拷贝数,执行分割和评估cnv特定的质量控制指标。PennCNV具有相对较小的偏倚,较小的变异性和检测更多的区域,同时保持与CRLMM/VanillaIce相似的估计假阳性率。更一般地说,我们主张软件开发人员需要提供关于评估和选择最佳设置的指导,以便为单个数据集获得最佳结果。在这样的指导存在之前,我们建议尝试多种算法,评估一致性/不一致性,然后考虑下游关联测试的区域联合。
Copy number data are routinely being extracted from genome-wide association study chips using a variety of software. We empirically evaluated and compared four freely-available software packages designed for Affymetrix SNP chips to estimate copy number: Affymetrix Power Tools (APT), Aroma.Affymetrix, PennCNV and CRLMM. Our evaluation used 1,418 GENOA samples that were genotyped on the Affymetrix Genome-Wide Human SNP Array 6.0. We compared bias and variance in the locus-level copy number data, the concordance amongst regions of copy number gains/deletions and the false-positive rate amongst deleted segments. APT had median locus-level copy numbers closest to a value of two, whereas PennCNV and Aroma.Affymetrix had the smallest variability associated with the median copy number. Of those evaluated, only PennCNV provides copy number specific quality-control metrics and identified 136 poor CNV samples. Regions of copy number variation (CNV) were detected using the hidden Markov models provided within PennCNV and CRLMM/VanillaIce. PennCNV detected more CNVs than CRLMM/VanillaIce; the median number of CNVs detected per sample was 39 and 30, respectively. PennCNV detected most of the regions that CRLMM/VanillaIce did as well as additional CNV regions. The median concordance between PennCNV and CRLMM/VanillaIce was 47.9% for duplications and 51.5% for deletions. The estimated false-positive rate associated with deletions was similar for PennCNV and CRLMM/VanillaIce. If the objective is to perform statistical tests on the locus-level copy number data, our empirical results suggest that PennCNV or Aroma.Affymetrix is optimal. If the objective is to perform statistical tests on the summarized segmented data then PennCNV would be preferred over CRLMM/VanillaIce. Specifically, PennCNV allows the analyst to estimate locus-level copy number, perform segmentation and evaluate CNV-specific quality-control metrics within a single software package. PennCNV has relatively small bias, small variability and detects more regions while maintaining a similar estimated false-positive rate as CRLMM/VanillaIce. More generally, we advocate that software developers need to provide guidance with respect to evaluating and choosing optimal settings in order to obtain optimal results for an individual dataset. Until such guidance exists, we recommend trying multiple algorithms, evaluating concordance/discordance and subsequently consider the union of regions for downstream association tests.
DOI: 10.2307/2987937
发表时间: 1983-01-01
期刊: JOURNAL OF THE ROYAL STATISTICAL SOCIETY SERIES D-THE STATISTICIAN
影响因子: --
作者:
ALTMAN, DG;BLAND, JM
通讯作者: BLAND, JM
DOI: 10.1093/bfgp/elp017
发表时间: 2009-09-01
期刊: Briefings in Functional Genomics & Proteomics
影响因子: --
作者:
Winchester, Laura;Yau, Christopher;Ragoussis, Jiannis
通讯作者: Ragoussis, Jiannis
DOI: 10.1093/nar/gkn556
发表时间: 2008-11
影响因子: 14.9
作者:
Diskin SJ;Li M;Hou C;Yang S;Glessner J;Hakonarson H;Bucan M;Maris JM;Wang K
通讯作者: Wang K
DOI: 10.1093/biostatistics/4.2.249
发表时间: 2003-04-01
期刊: BIOSTATISTICS
影响因子: 2.1
作者:
Irizarry, RA;Hobbs, B;Speed, TP
通讯作者: Speed, TP
DOI: 10.1093/bioinformatics/btn016
发表时间: 2008-03-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Bengtsson, H.;Irizarry, R.;Speed, T. P.
通讯作者: Speed, T. P.