A community-based resource for automatic exome variant-calling and annotation in Mendelian disorders.

A community-based resource for automatic exome variant-calling and annotation in Mendelian disorders.
复制标题

一种基于社区的资源,用于Mendelian疾病中的自动外显子变体呼叫和注释。

DOI:
10.1186/1471-2164-15-s3-s5
复制
发表时间:
2014
期刊:
影响因子:
4.4
通讯作者:
di Bernardo D
di Bernardo D
中科院分区:
生物学2区
文献类型:
--
作者:
Mutarelli M;Marwah V;Rispoli R;Carrella D;Dharmalingam G;Oliva G;di Bernardo D

文献摘要

被引文献

相似文献

孟德尔疾病主要由基因DNA序列的单一突变引起,导致具有病理后果的表型。患者的全外显子组测序可以是一种具有成本效益的替代标准遗传筛查,以发现遗传性疾病的致病突变,特别是当病例数量有限时。分析外显子组测序数据需要特定的专业知识、高计算资源和参考变异数据库来识别致病变异。我们开发了一个从孟德尔疾病患者中收集的变异数据库,由于相关的外显子组测序管道,该数据库可以自动填充。该管道能够自动识别、注释和存储数据库中的插入、删除和突变。该资源可在网上免费获得http://exome.tigem.it。外显子组测序流水线使用最先进的软件工具自动化分析工作流程(质量控制和读取修剪,参考基因组图谱,比对后处理,变异调用和注释)。外显子组测序流水线被设计为在一个计算集群上运行,以便同时分析多个样本。管道不仅用标准的变异注释(例如,一般人群中的等位基因频率,对基因产物活性的预测影响等)注释检测到的变异,而且更重要的是,在数据库中逐步收集样本中的等位基因频率,并按孟德尔紊乱分层。我们的目标是为遗传病界提供一个资源,通过标准和统一的分析管道自动分析全外显子组测序样本,从而按疾病收集变异等位基因频率。这种资源可能成为一个有价值的工具,通过改进对假定的患者特异性致病或表型相关变异的选择,帮助解剖疾病表型背后的基因型。
Mendelian disorders are mostly caused by single mutations in the DNA sequence of a gene, leading to a phenotype with pathologic consequences. Whole Exome Sequencing of patients can be a cost-effective alternative to standard genetic screenings to find causative mutations of genetic diseases, especially when the number of cases is limited. Analyzing exome sequencing data requires specific expertise, high computational resources and a reference variant database to identify pathogenic variants. We developed a database of variations collected from patients with Mendelian disorders, which is automatically populated thanks to an associated exome-sequencing pipeline. The pipeline is able to automatically identify, annotate and store insertions, deletions and mutations in the database. The resource is freely available online http://exome.tigem.it. The exome sequencing pipeline automates the analysis workflow (quality control and read trimming, mapping on reference genome, post-alignment processing, variation calling and annotation) using state-of-the-art software tools. The exome-sequencing pipeline has been designed to run on a computing cluster in order to analyse several samples simultaneously. The detected variants are annotated by the pipeline not only with the standard variant annotations (e.g. allele frequency in the general population, the predicted effect on gene product activity, etc.) but, more importantly, with allele frequencies across samples progressively collected in the database itself, stratified by Mendelian disorder. We aim at providing a resource for the genetic disease community to automatically analyse whole exome-sequencing samples with a standard and uniform analysis pipeline, thus collecting variant allele frequencies by disorder. This resource may become a valuable tool to help dissecting the genotype underlying the disease phenotype through an improved selection of putative patient-specific causative or phenotype-associated variations.