Improving sequence variant descriptions in mutation Databases and literature using the mutalyzer sequence variation nomenclature checker

Improving sequence variant descriptions in mutation Databases and literature using the mutalyzer sequence variation nomenclature checker
复制标题

DOI:
10.1002/humu.20654
复制
发表时间:
2008-01-01
期刊:
影响因子:
3.9
通讯作者:
Taschner, Peter E. M.
Taschner, Peter E. M.
中科院分区:
医学2区
文献类型:
--
作者:
Wildeman, Martin;van Ophuizen, Ernest;Taschner, Peter E. M.

文献摘要

被引文献

相似文献

明确和正确的序列变异描述是至关重要的,因为错误和不确定性可能导致临床诊断中的不期望错误。我们开发了突变分析仪(Mutalyzer)序列变异命名检查器(www.lovd.nl/mutalyzer;最后访问2007年9月13日),用于使用任何生物的参考序列自动分析和纠正序列变异描述。Mutalyzer处理大多数变异类型:替换、删除、复制、插入、indel和剪接,位点变化遵循人类基因组变异学会(HGVS)的当前建议。输入是GenBank的加入号或上传的参考序列文件(GenBank格式,带有用户修改的注释),HGNC基因符号和变体(单个或批处理文件)。Mutalyzer在DNA水平上产生变异描述,在所有注释转录本的水平上产生变异描述,在蛋白质水平上产生推断结果。为了验证Mutalyzer的性能并研究位点特异性突变数据库(lsdb)中的序列变异描述质量,对PAH、BIC BRCA2和HbVar数据库中的11,000多个变异进行了分析,结果显示,分别有87%、25%和38%的变异没有错误,并遵循了建议。BIC和HbVar的低识别率(分别为38%和51%)是由于缺乏注释良好的基因组参考序列(HbVar)或不遵守指南(BRCA2)。Mutalyzer提供了注释良好的基因组参考序列,对于新发现的序列变异描述和现有的LSDB数据的管理非常有效。Mutalyzer将连接到Leiden开源变异数据库(LOVD) (www.LOVD.nl;最后一次访问是2007年9月13日),并且是序列变异效应预测包的第一个模块。
Unambiguous and correct sequence variant descriptions are of utmost importance, not in the least since mistakes and uncertainties may lead to undesired errors in clinical diagnosis. We developed the Mutation Analyzer (Mutalyzer) sequence variation nomenclature checker (www.lovd.nl/mutalyzer; last accessed 13 September 2007) for automated analysis and correction of sequence variant descriptions using reference sequences from any organism. Mutalyzer handles most variation types: substitution, deletion, duplication, insertion, indel, and splice,site changes following current recommendations of the Human Genome Variation Society (HGVS). Input is a GenBank accession number or an uploaded reference sequence file in GenBank format with user-modified annotation, an HGNC gene symbol, and the variant (single or in a batch file). Mutalyzer generates variant descriptions at DNA level, the level of all annotated transcripts and the deduced outcome at protein level. To validate Mutalyzer's performance and to investigate the sequence variant description quality in locus-specific mutation databases (LSDBs), more than 11,000 variants in the PAH, BIC BRCA2, and HbVar databases were analyzed, showing that 87%, 25%, and 38%, respectively, were error-free and following the recommendations. Low recognition rates in BIC and HbVar (38% and 51%, respectively) were due to lack of a well-annotated genomic reference sequence (HbVar) or noncompliance to the guidelines (BRCA2). Provided with well-annotated genomic reference sequences, Mutalyzer is very effective for the curation of newly discovered sequence variation descriptions and existing LSDB, data. Mutalyzer will be linked to the Leiden Open source Variation Database (LOVD) (www.LOVD.nl; last accessed 13 September 2007) and is the first module of a sequence variant effect prediction package.