VAAST 2.0: improved variant classification and disease-gene identification using a conservation-controlled amino acid substitution matrix.

VAAST 2.0: improved variant classification and disease-gene identification using a conservation-controlled amino acid substitution matrix.
复制标题

DOI:
10.1002/gepi.21743
复制
发表时间:
2013-09
影响因子:
2.1
通讯作者:
Yandell, Mark
Yandell, Mark
中科院分区:
医学4区
文献类型:
--
作者:
Hu, Hao;Huff, Chad D.;Moore, Barry;Flygare, Steven;Reese, Martin G.;Yandell, Mark

文献摘要

参考文献

被引文献

相似文献

人们普遍认为,需要改进算法,以支持个人基因组数据中的变异优先级排序和疾病基因识别。我们之前介绍了变体注释,分析和搜索工具(VAAST),它采用了一个聚合的变体关联测试,结合了氨基酸取代(AAS)和等位基因频率。在这里,我们描述和基准VAAST 2.0,它使用了一种新的保护控制的AAS矩阵(CASM),将有关系统发育保护的信息。我们表明,CASM方法提高了VAAST的变体优先级排序的准确性相比,其以前的实现,相比SIFT,PolyPhen-2,和MutationTaster。我们还表明,VAAST 2.0优于KBAC,WSS,SKAT和可变阈值(VT)使用已发表的病例对照数据集克罗恩病(NOD 2),高脂血症(LPL)和乳腺癌(CHEK 2)。VAAST 2.0还提高了在广泛的等位基因频率、人群可归因疾病风险和等位基因异质性等影响其他聚集性变异相关性检验准确性的因素中对模拟数据集的搜索准确性。我们还证明,虽然大多数聚集性变异关联测试是针对常见的遗传疾病设计的,但这些测试可以很容易地被采用为罕见的孟德尔疾病基因发现者,具有简单的按性别显著性排名协议,并且性能与最先进的过滤方法相比非常有利。后者,尽管他们的普及,有次优性能,特别是随着案例样本量的增加。
The need for improved algorithmic support for variant prioritization and disease-gene identification in personal genomes data is widely acknowledged. We previously presented the Variant Annotation, Analysis, and Search Tool (VAAST), which employs an aggregative variant association test that combines both amino acid substitution (AAS) and allele frequencies. Here we describe and benchmark VAAST 2.0, which uses a novel conservation-controlled AAS matrix (CASM), to incorporate information about phylogenetic conservation. We show that the CASM approach improves VAAST’s variant prioritization accuracy compared to its previous implementation, and compared to SIFT, PolyPhen-2, and MutationTaster. We also show that VAAST 2.0 outperforms KBAC, WSS, SKAT, and variable threshold (VT) using published case-control datasets for Crohn disease (NOD2), hypertriglyceridemia (LPL), and breast cancer (CHEK2). VAAST 2.0 also improves search accuracy on simulated datasets across a wide range of allele frequencies, population-attributable disease risks, and allelic heterogeneity, factors that compromise the accuracies of other aggregative variant association tests. We also demonstrate that, although most aggregative variant association tests are designed for common genetic diseases, these tests can be easily adopted as rare Mendelian disease-gene finders with a simple ranking-by-statistical-significance protocol, and the performance compares very favorably to state-of-art filtering approaches. The latter, despite their popularity, have suboptimal performance especially with the increasing case sample size.
DOI: 10.1186/bcr2810
发表时间: 2011-01-18
期刊: Breast cancer research : BCR
影响因子: --
作者:
Le Calvez-Kelm F;Lesueur F;Damiola F;Vallée M;Voegele C;Babikyan D;Durand G;Forey N;McKay-Chopin S;Robinot N;Nguyen-Dumont T;Thomas A;Byrnes GB;Breast Cancer Family Registry;Hopper JL;Southey MC;Andrulis IL;John EM;Tavtigian SV
通讯作者: Tavtigian SV
DOI: 10.1086/521032
发表时间: 2007-11-01
影响因子: 9.8
作者:
Easton, Douglas F.;Deffenbaugh, Amie M.;Goldgar, David E.
通讯作者: Goldgar, David E.
DOI: 10.1371/journal.pgen.1001156
发表时间: 2010-10-14
期刊: PLoS genetics
影响因子: 4.5
作者:
Liu DJ;Leal SM
通讯作者: Leal SM
DOI: 10.1038/ng.628
发表时间: 2010-08
期刊: Nature genetics
影响因子: 30.8
作者:
通讯作者: --
DOI: 10.1016/j.ajhg.2008.06.024
发表时间: 2008-09-12
影响因子: 9.8
作者:
Li, Bingshan;Leal, Suzanne M.
通讯作者: Leal, Suzanne M.