Rich annotation of DNA sequencing variants by leveraging the Ensembl Variant Effect Predictor with plugins

Rich annotation of DNA sequencing variants by leveraging the Ensembl Variant Effect Predictor with plugins
复制标题

DOI:
10.1093/bib/bbu008
复制
发表时间:
2015-03-01
影响因子:
9.5
通讯作者:
Nelson, Stanley F.
Nelson, Stanley F.
中科院分区:
生物学2区
文献类型:
--
作者:
Yourshaw, Michael;Taylor, S. Paige;Nelson, Stanley F.

文献摘要

被引文献

相似文献

高通量DNA测序已成为发现可能导致疾病或影响表型的基因组变异的支柱。下一代测序管道通常识别每个样品中的数千个变体。一个特别的挑战是以一种对数据下游消费者(例如临床测序中心或研究人员)有用的方式注释每个变体。这些用户可能需要将所有数据存储和分析保留在安全的本地服务器上,以保护患者的机密性或知识产权,可能有独特的和不断变化的需求,以利用各种注释数据集,并且可能更喜欢不依赖于他们无法控制的闭源应用程序。在这里,我们描述了可扩展的方法,用于使用Ensembl变体效应预测器的插件功能来丰富其基本的变体注释集,其中包含有关基因、功能、保守性、表达、疾病、途径和蛋白质结构的额外数据,并描述了一个可扩展的框架,用于轻松添加额外的自定义数据集。
High-throughput DNA sequencing has become a mainstay for the discovery of genomic variants that may cause disease or affect phenotype. A next-generation sequencing pipeline typically identifies thousands of variants in each sample. A particular challenge is the annotation of each variant in a way that is useful to downstream consumers of the data, such as clinical sequencing centers or researchers. These users may require that all data storage and analysis remain on secure local servers to protect patient confidentiality or intellectual property, may have unique and changing needs to draw on a variety of annotation data sets and may prefer not to rely on closed-source applications beyond their control. Here we describe scalable methods for using the plugin capability of the Ensembl Variant Effect Predictor to enrich its basic set of variant annotations with additional data on genes, function, conservation, expression, diseases, pathways and protein structure, and describe an extensible framework for easily adding additional custom data sets.