EnsembleCNV: an ensemble machine learning algorithm to identify and genotype copy number variation using SNP array data.

EnsembleCNV: an ensemble machine learning algorithm to identify and genotype copy number variation using SNP array data.
复制标题

EnsembleCNV:一种集成机器学习算法,用于使用 SNP 阵列数据识别拷贝数变异并对其进行基因分型。

DOI:
10.1093/nar/gkz068
复制
发表时间:
2019
影响因子:
14.9
通讯作者:
Hao,Ke
Hao,Ke
中科院分区:
生物学2区
文献类型:
--
作者:
Zhang,Zhongyang;Cheng,Haoxiang;Hong,Xiumei;DiNarzo,AntonioF;Franzen,Oscar;Peng,Shouneng;Ruusalepp,Arno;Kovacic,JasonC;Bjorkegren,JohanLM;Wang,Xiaobin;Hao,Ke

文献摘要

相似文献

疾病/性状和拷贝数变异(CNVs)之间的关联尚未在全基因组关联研究(GWAS)中进行系统研究,这主要是由于缺乏可靠和准确的CNV基因分型工具。在这里,我们提出了一种新的集成学习框架,ensembleCNV,检测和基因型CNVs使用单核苷酸多态性(SNP)阵列数据。EnsembleCNV(a)在原始数据水平上鉴定并消除批次效应;(B)通过启发式算法将来自多个具有互补强度的现有识别者的单个CNV识别组装成CNV区域(CNVR);(c)用局部似然模型对每个CNVR进行重新基因型化,所述局部似然模型通过跨多个CNVR的全局信息进行调整;(d)通过拷贝数强度中的局部相关性结构来细化CNVR边界;(e)提供直接的CNV基因分型,伴随着置信度评分,可直接用于下游质量控制和关联分析。在两个大型数据集上进行基准测试,ensembleCNV优于竞争方法,并实现了高调用率(93.3%)和再现性(98.6%),同时通过捕获1000个基因组计划中记录的85%的常见CNV来实现高灵敏度。鉴于这种CNV调用率和准确性,这是可比的SNP基因分型,我们建议ensembleCNV具有显着的承诺进行全基因组CNV关联研究和调查CNV如何易感于人类疾病。
The associations between diseases/traits and copy number variants (CNVs) have not been systematically investigated in genome-wide association studies (GWASs), primarily due to a lack of robust and accurate tools for CNV genotyping. Herein, we propose a novel ensemble learning framework, ensembleCNV, to detect and genotype CNVs using single nucleotide polymorphism (SNP) array data. EnsembleCNV (a) identifies and eliminates batch effects at raw data level; (b) assembles individual CNV calls into CNV regions (CNVRs) from multiple existing callers with complementary strengths by a heuristic algorithm; (c) re-genotypes each CNVR with local likelihood model adjusted by global information across multiple CNVRs; (d) refines CNVR boundaries by local correlation structure in copy number intensities; (e) provides direct CNV genotyping accompanied with confidence score, directly accessible for downstream quality control and association analysis. Benchmarked on two large datasets, ensembleCNV outperformed competing methods and achieved a high call rate (93.3%) and reproducibility (98.6%), while concurrently achieving high sensitivity by capturing 85% of common CNVs documented in the 1000 Genomes Project. Given this CNV call rate and accuracy, which are comparable to SNP genotyping, we suggest ensembleCNV holds significant promise for performing genome-wide CNV association studies and investigating how CNVs predispose to human diseases.