WGSA: an annotation pipeline for human genome sequencing studies.

WGSA: an annotation pipeline for human genome sequencing studies.
复制标题

DOI:
10.1136/jmedgenet-2015-103423
复制
发表时间:
2016-02
影响因子:
4
通讯作者:
Boerwinkle E
Boerwinkle E
中科院分区:
医学1区
文献类型:
--
作者:
Liu X;White S;Peng B;Johnson AD;Brody JA;Li AH;Huang Z;Carroll A;Wei P;Gibbs R;Klein RJ;Boerwinkle E

文献摘要

参考文献

被引文献

相似文献

DNA测序技术在提高通量和质量以及降低成本方面继续取得进展。当我们从全外显子组捕获测序过渡到全基因组测序(WGS)时,我们将机器生成的变异识别(包括单核苷酸变异(SNV)和插入-缺失变异(indels))转化为人类可解释的知识的能力远远落后于获得大量变异的能力。为了帮助缩小这一差距,我们在这里介绍了WGSA(WGS注释器),这是一种用于人类基因组测序研究的功能注释管道,可在Amazon Compute Cloud上开箱即用,并可在(https://sites. Google. com/site/jpopgen/wgsa/)。功能注释是WGS分析的关键步骤。在某种程度上,注释帮助分析人员筛选出特别感兴趣的元素子集(例如,细胞类型特异性增强子),在另一种方式下,注释帮助研究人员提高识别表型相关基因座的能力(例如,使用功能预测评分作为权重的关联测试)并解释潜在的有趣发现。目前,有几种流行的基于基因模型的注释工具,包括ANNOVAR、1 SnpEff 2和Ensembl变体效应预测器(VEP)。3这些可以注释来自一系列物种的各种蛋白质编码和非编码基因模型。从业者中众所周知的是,不同的数据库(例如,RefSeq 4和Ensembl 5)对相同的基因使用不同的模型。即使实现了相同的基因结构,不同注释工具对给定变体的预测结果也可能不一致。6因此,有人建议从多个数据库的工具中获得注释,以更完整地解释WGS中发现的变体。6编码和非编码变异的注释包括与功能性、保守性、群体等位基因频率和疾病相关注释有关的评分,即在全基因组关联分析中鉴定的已知致病变异和疾病相关变异。最近的大规模表观基因组学项目提供了丰富的细胞特异性调控元件的数据集。不幸的是,目前很少有工具可用于整合所有这些功能注释资源,并提供一个方便和有效的管道来注释WGS研究中发现的数百万个变体。为了方便WGS的功能注释步骤,我们开发了WGSA。目前WGSA支持本地SNV和indel的注释,而无需远程数据库请求,使其能够扩展到大型WGS研究。WGSA管道概述见图1。工作组所载资源(及其参考资料)的完整清单可在在线补充表S1中找到。对于基于基因模型的注释,WGSA整合了三种注释工具(ANNOVAR,SnpEff和VEP)与两个数据库(RefSeq和Ensembl)的输出,并提供了六个注释结果的变体结果总结。为了进一步加快大规模WGS研究的过程,我们基于人类参考hg 19非N碱基预先计算了所有潜在人类SNV(共8 584 031 106个)的注释,并将其用作本地数据库。对于以SNV为中心的资源,WGSA整合了五个功能预测评分,八个保守评分,来自四个大规模测序研究的等位基因频率,四个疾病相关数据库中的变体等(图1和在线补充表S1)。对于以调控区域为中心的资源,WGSA包括细胞类型特异性转录因子...
DNA sequencing technologies continue to make progress in increased throughput and quality, and decreased cost. As we transition from whole exome capture sequencing to whole genome sequencing (WGS), our ability to convert machine-generated variant calls, including single nucleotide variant (SNV) and insertion-deletion variants (indels), into human-interpretable knowledge has lagged far behind the ability to obtain enormous amounts of variants. To help narrow this gap, here we present WGSA (WGS annotator), a functional annotation pipeline for human genome sequencing studies, which is runnable out of the box on the Amazon Compute Cloud and is freely downloadable at (https://sites. google. com/site/jpopgen/wgsa/). Functional annotation is a key step in WGS analysis. In one way, annotation helps the analyst filter to a subset of elements of particular interest (eg, cell type specific enhancers), in another way annotation helps the investigators to increase the power of identifying phenotypeassociated loci (eg, association test using functional prediction score as a weight) and interpret potentially interesting findings. Currently, there are several popular gene model based annotation tools, including ANNOVAR, 1 SnpEff 2 and the Ensembl Variant Effect Predictor (VEP). 3 These can annotate a variety of protein coding and non-coding gene models from a range of species. It is well known among practitioners that different databases (eg, RefSeq 4 and Ensembl 5) use different models for the same gene. Even when the same gene structure is implemented, predicted consequences of a given variant from different annotation tools may not be in agreement. 6 Therefore, it has been suggested to obtain annotation from tools across multiple databases for a more complete interpretation of the variants discovered in WGS. 6 Annotations of coding and non-coding variants include scores pertaining to functionality, conservation, population allele frequencies and disease-related annotations, that is, known disease-causing variants and disease-associated variants identified in genome-wide association analyses. Recent large-scale epigenomics projects provide rich data sets of cell-specific regulatory elements. Unfortunately, there are currently few tools available to integrate all those functional annotation resources and provide a convenient and efficient pipeline for annotating millions of variants discovered in a WGS study. To facilitate the functional annotation step of WGS, we developed WGSA.Currently WGSA supports the annotation of SNVs and indels locally without remote database requests, allowing it to scale up for large WGS studies. The overview of the WGSA pipeline is presented in figure 1. The complete list of the resources (and their references) contained in WGSA can be found in online supplementary table S1. For gene-model based annotation, WGSA integrates the outputs from three annotation tools (ANNOVAR, SnpEff and VEP) versus two databases (RefSeq and Ensembl), and provides a summary of variant consequences from the six annotation results. To further speed up the process for large-scale WGS studies, we have precomputed annotations for all potential human SNVs (a total of 8 584 031 106) based on human reference hg19 non-N bases and use it as a local database. For SNV-centric resources, WGSA integrates five functional prediction scores, eight conservation scores, allele frequencies from four large-scale sequencing studies, variants in four disease-related databases, among others (figure 1 and online supplementary table S1). For regulatory region-centric resources, WGSA includes cell type specific transcription factor …
DOI: 10.1093/bioinformatics/btq330
发表时间: 2010-08-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
McLaren, William;Pritchard, Bethan;Cunningham, Fiona
通讯作者: Cunningham, Fiona
DOI: 10.4161/fly.19695
发表时间: 2012-04-01
期刊: FLY
影响因子: 1.2
作者:
Cingolani, Pablo;Platts, Adrian;Ruden, Douglas M.
通讯作者: Ruden, Douglas M.
DOI: 10.1186/gm543
发表时间: 2014
期刊: Genome medicine
影响因子: 12.3
作者:
McCarthy DJ;Humburg P;Kanapin A;Rivas MA;Gaulton K;Cazier JB;Donnelly P
通讯作者: Donnelly P
DOI: 10.1093/nar/gku1010
发表时间: 2015-01
影响因子: 14.9
作者:
Cunningham F;Amode MR;Barrell D;Beal K;Billis K;Brent S;Carvalho-Silva D;Clapham P;Coates G;Fitzgerald S;Gil L;Girón CG;Gordon L;Hourlier T;Hunt SE;Janacek SH;Johnson N;Juettemann T;Kähäri AK;Keenan S;Martin FJ;Maurel T;McLaren W;Murphy DN;Nag R;Overduin B;Parker A;Patricio M;Perry E;Pignatelli M;Riat HS;Sheppard D;Taylor K;Thormann A;Vullo A;Wilder SP;Zadissa A;Aken BL;Birney E;Harrow J;Kinsella R;Muffato M;Ruffier M;Searle SM;Spudich G;Trevanion SJ;Yates A;Zerbino DR;Flicek P
通讯作者: Flicek P
DOI: 10.1093/nar/gkq603
发表时间: 2010-09
影响因子: 14.9
作者:
Wang K;Li M;Hakonarson H
通讯作者: Hakonarson H