Automated validation of genetic variants from large databases: ensuring that variant references refer to the same genomic locations

Automated validation of genetic variants from large databases: ensuring that variant references refer to the same genomic locations
复制标题

DOI:
10.1093/bioinformatics/btr029
复制
发表时间:
2011-03-15
期刊:
影响因子:
5.8
通讯作者:
Kohane, Isaac S.
Kohane, Isaac S.
中科院分区:
生物学3区
文献类型:
--
作者:
Tong, Mark Y.;Cassa, Christopher A.;Kohane, Isaac S.

文献摘要

被引文献

相似文献

总结:基因组变异的准确注释对于实现科学合理且医学相关的全基因组临床解释是必要的。许多疾病关联,特别是那些在人类基因组计划完成之前报道的,由于与我们目前的基因组坐标,命名和基因结构标准可能不一致,因此适用性有限。为了验证并将来自医学遗传学文献的变体与每个变体的明确参考联系起来,我们开发了一个软件管道,并回顾了来自在线孟德尔遗传人类(OMIM),人类基因突变数据库(HGMD)和dbSNP的68641个单氨基酸突变。未解决的突变注释的频率在数据库中变化很大,范围从4%到23%。产生了未解决的突变的主要原因的分类。
Summary: Accurate annotations of genomic variants are necessary to achieve full-genome clinical interpretations that are scientifically sound and medically relevant. Many disease associations, especially those reported before the completion of the HGP, are limited in applicability because of potential inconsistencies with our current standards for genomic coordinates, nomenclature and gene structure. In an effort to validate and link variants from the medical genetics literature to an unambiguous reference for each variant, we developed a software pipeline and reviewed 68 641 single amino acid mutations from Online Mendelian Inheritance in Man (OMIM), Human Gene Mutation Database (HGMD) and dbSNP. The frequency of unresolved mutation annotations varied widely among the databases, ranging from 4 to 23%. A taxonomy of primary causes for unresolved mutations was produced.