Software comparison for evaluating genomic copy number variation for Affymetrix 6.0 SNP array platform.
Software comparison for evaluating genomic copy number variation for Affymetrix 6.0 SNP array platform.
复制标题
DOI:
10.1186/1471-2105-12-220
复制
发表时间:
2011-05-31
影响因子:
3
通讯作者:
de Andrade M
中科院分区:
文献类型:
--
作者:
Eckel-Passow JE;Atkinson EJ;Maharjan S;Kardia SL;de Andrade M
Copy number data are routinely being extracted from genome-wide association study chips using a variety of software. We empirically evaluated and compared four freely-available software packages designed for Affymetrix SNP chips to estimate copy number: Affymetrix Power Tools (APT), Aroma.Affymetrix, PennCNV and CRLMM. Our evaluation used 1,418 GENOA samples that were genotyped on the Affymetrix Genome-Wide Human SNP Array 6.0. We compared bias and variance in the locus-level copy number data, the concordance amongst regions of copy number gains/deletions and the false-positive rate amongst deleted segments. APT had median locus-level copy numbers closest to a value of two, whereas PennCNV and Aroma.Affymetrix had the smallest variability associated with the median copy number. Of those evaluated, only PennCNV provides copy number specific quality-control metrics and identified 136 poor CNV samples. Regions of copy number variation (CNV) were detected using the hidden Markov models provided within PennCNV and CRLMM/VanillaIce. PennCNV detected more CNVs than CRLMM/VanillaIce; the median number of CNVs detected per sample was 39 and 30, respectively. PennCNV detected most of the regions that CRLMM/VanillaIce did as well as additional CNV regions. The median concordance between PennCNV and CRLMM/VanillaIce was 47.9% for duplications and 51.5% for deletions. The estimated false-positive rate associated with deletions was similar for PennCNV and CRLMM/VanillaIce. If the objective is to perform statistical tests on the locus-level copy number data, our empirical results suggest that PennCNV or Aroma.Affymetrix is optimal. If the objective is to perform statistical tests on the summarized segmented data then PennCNV would be preferred over CRLMM/VanillaIce. Specifically, PennCNV allows the analyst to estimate locus-level copy number, perform segmentation and evaluate CNV-specific quality-control metrics within a single software package. PennCNV has relatively small bias, small variability and detects more regions while maintaining a similar estimated false-positive rate as CRLMM/VanillaIce. More generally, we advocate that software developers need to provide guidance with respect to evaluating and choosing optimal settings in order to obtain optimal results for an individual dataset. Until such guidance exists, we recommend trying multiple algorithms, evaluating concordance/discordance and subsequently consider the union of regions for downstream association tests.
登录
查看更多内容
DOI:
10.2307/2987937
发表时间:
1983-01-01
期刊:
JOURNAL OF THE ROYAL STATISTICAL SOCIETY SERIES D-THE STATISTICIAN
影响因子:
--
作者:
ALTMAN, DG;BLAND, JM
通讯作者:
BLAND, JM
DOI:
10.1093/bfgp/elp017
发表时间:
2009-09-01
期刊:
Briefings in Functional Genomics & Proteomics
影响因子:
--
作者:
Winchester, Laura;Yau, Christopher;Ragoussis, Jiannis
通讯作者:
Ragoussis, Jiannis
影响因子:
14.9
作者:
Diskin SJ;Li M;Hou C;Yang S;Glessner J;Hakonarson H;Bucan M;Maris JM;Wang K
通讯作者:
Wang K
影响因子:
2.1
作者:
Irizarry, RA;Hobbs, B;Speed, TP
通讯作者:
Speed, TP
影响因子:
5.8
作者:
Bengtsson, H.;Irizarry, R.;Speed, T. P.
通讯作者:
Speed, T. P.