Comparative analysis of copy number variation detection methods and database construction

Comparative analysis of copy number variation detection methods and database construction
复制标题

DOI:
10.1186/1471-2156-12-29
复制
发表时间:
2011-03-07
期刊:
影响因子:
2.9
通讯作者:
Tokunaga, Katsushi
Tokunaga, Katsushi
中科院分区:
生物学3区
文献类型:
--
作者:
Koike, Asako;Nishida, Nao;Tokunaga, Katsushi

文献摘要

被引文献

相似文献

背景:基于阵列的拷贝数变异检测(CNVs)被广泛用于识别疾病特异性遗传变异。然而,CNV检测的准确性是不够的,结果取决于所使用的检测程序和它们的参数。在本研究中,我们利用HapMap数据和其他实验数据,从Affymetrix平台上的性能角度评估了五个广泛使用的CNV检测程序,Birdsuite(主要由birdeye和Canary模块组成)、birdeye (Birdsuite的一部分)、PennCNV、CGHseg和DNAcopy。此外,我们使用在HapMap数据中表现最佳的参数确定了180名健康日本人的CNVs,并研究了他们的特征。结果:考虑到相同个体的高再现率和低孟德尔不一致性,基于隐马尔可夫模型的程序PennCNV和birdeye (Birdsuite的一部分)或Birdsuite的检测性能优于其他程序。此外,当考虑到与其他实验结果的重叠率时,Birdsuite从灵敏度的角度显示出最佳性能,但预计会包括许多假阴性和一些假阳性。180例健康日本人的结果表明,在多个体间普遍检测到的CNVs的起始和结束区域,包含重复序列的比例(不仅是片段重复序列,还有长间隔核元件(long interspersed nuclear element, LINE)序列)比随机选择的区域要高,并且基于灵长类动物的保守评分比随机选择的区域要低。在HapMap数据和其他实验数据中也观察到类似的趋势。结论:我们的研究结果表明,不仅片段重复序列,而且穿插重复序列,特别是LINE序列,都与CNV有关,特别是在常见的CNV形成中。检测到的CNV存储在“日本综合数据库项目”新建的CNV存储库数据库中,供研究人员共享数据。http://gwas.lifesciencedb.jp/cgi-bin/cnvdb/cnv_top.cgi。
Background: Array-based detection of copy number variations (CNVs) is widely used for identifying disease-specific genetic variations. However, the accuracy of CNV detection is not sufficient and results differ depending on the detection programs used and their parameters. In this study, we evaluated five widely used CNV detection programs, Birdsuite (mainly consisting of the Birdseye and Canary modules), Birdseye ( part of Birdsuite), PennCNV, CGHseg, and DNAcopy from the viewpoint of performance on the Affymetrix platform using HapMap data and other experimental data. Furthermore, we identified CNVs of 180 healthy Japanese individuals using parameters that showed the best performance in the HapMap data and investigated their characteristics.Results: The results indicate that Hidden Markov model-based programs PennCNV and Birdseye (part of Birdsuite), or Birdsuite show better detection performance than other programs when the high reproducibility rates of the same individuals and the low Mendelian inconsistencies are considered. Furthermore, when rates of overlap with other experimental results were taken into account, Birdsuite showed the best performance from the view point of sensitivity but was expected to include many false negatives and some false positives. The results of 180 healthy Japanese demonstrate that the ratio containing repeat sequences, not only segmental repeats but also long interspersed nuclear element (LINE) sequences both in the start and end regions of the CNVs, is higher in CNVs that are commonly detected among multiple individuals than that in randomly selected regions, and the conservation score based on primates is lower in these regions than in randomly selected regions. Similar tendencies were observed in HapMap data and other experimental data.Conclusions: Our results suggest that not only segmental repeats but also interspersed repeats, especially LINE sequences, are deeply involved in CNVs, particularly in common CNV formations. The detected CNVs are stored in the CNV repository database newly constructed by the "Japanese integrated database project" for sharing data among researchers. http://gwas.lifesciencedb.jp/cgi-bin/cnvdb/cnv_top.cgi.