WGSA: an annotation pipeline for human genome sequencing studies.
WGSA: an annotation pipeline for human genome sequencing studies.
复制标题
DOI:
10.1136/jmedgenet-2015-103423
复制
发表时间:
2016-02
影响因子:
4
通讯作者:
Boerwinkle E
中科院分区:
文献类型:
--
作者:
Liu X;White S;Peng B;Johnson AD;Brody JA;Li AH;Huang Z;Carroll A;Wei P;Gibbs R;Klein RJ;Boerwinkle E
DNA sequencing technologies continue to make progress in increased throughput and quality, and decreased cost. As we transition from whole exome capture sequencing to whole genome sequencing (WGS), our ability to convert machine-generated variant calls, including single nucleotide variant (SNV) and insertion-deletion variants (indels), into human-interpretable knowledge has lagged far behind the ability to obtain enormous amounts of variants. To help narrow this gap, here we present WGSA (WGS annotator), a functional annotation pipeline for human genome sequencing studies, which is runnable out of the box on the Amazon Compute Cloud and is freely downloadable at (https://sites. google. com/site/jpopgen/wgsa/). Functional annotation is a key step in WGS analysis. In one way, annotation helps the analyst filter to a subset of elements of particular interest (eg, cell type specific enhancers), in another way annotation helps the investigators to increase the power of identifying phenotypeassociated loci (eg, association test using functional prediction score as a weight) and interpret potentially interesting findings. Currently, there are several popular gene model based annotation tools, including ANNOVAR, 1 SnpEff 2 and the Ensembl Variant Effect Predictor (VEP). 3 These can annotate a variety of protein coding and non-coding gene models from a range of species. It is well known among practitioners that different databases (eg, RefSeq 4 and Ensembl 5) use different models for the same gene. Even when the same gene structure is implemented, predicted consequences of a given variant from different annotation tools may not be in agreement. 6 Therefore, it has been suggested to obtain annotation from tools across multiple databases for a more complete interpretation of the variants discovered in WGS. 6 Annotations of coding and non-coding variants include scores pertaining to functionality, conservation, population allele frequencies and disease-related annotations, that is, known disease-causing variants and disease-associated variants identified in genome-wide association analyses. Recent large-scale epigenomics projects provide rich data sets of cell-specific regulatory elements. Unfortunately, there are currently few tools available to integrate all those functional annotation resources and provide a convenient and efficient pipeline for annotating millions of variants discovered in a WGS study. To facilitate the functional annotation step of WGS, we developed WGSA.Currently WGSA supports the annotation of SNVs and indels locally without remote database requests, allowing it to scale up for large WGS studies. The overview of the WGSA pipeline is presented in figure 1. The complete list of the resources (and their references) contained in WGSA can be found in online supplementary table S1. For gene-model based annotation, WGSA integrates the outputs from three annotation tools (ANNOVAR, SnpEff and VEP) versus two databases (RefSeq and Ensembl), and provides a summary of variant consequences from the six annotation results. To further speed up the process for large-scale WGS studies, we have precomputed annotations for all potential human SNVs (a total of 8 584 031 106) based on human reference hg19 non-N bases and use it as a local database. For SNV-centric resources, WGSA integrates five functional prediction scores, eight conservation scores, allele frequencies from four large-scale sequencing studies, variants in four disease-related databases, among others (figure 1 and online supplementary table S1). For regulatory region-centric resources, WGSA includes cell type specific transcription factor …
登录
查看更多内容
影响因子:
5.8
作者:
McLaren, William;Pritchard, Bethan;Cunningham, Fiona
通讯作者:
Cunningham, Fiona
影响因子:
1.2
作者:
Cingolani, Pablo;Platts, Adrian;Ruden, Douglas M.
通讯作者:
Ruden, Douglas M.
影响因子:
12.3
作者:
McCarthy DJ;Humburg P;Kanapin A;Rivas MA;Gaulton K;Cazier JB;Donnelly P
通讯作者:
Donnelly P
影响因子:
14.9
作者:
Cunningham F;Amode MR;Barrell D;Beal K;Billis K;Brent S;Carvalho-Silva D;Clapham P;Coates G;Fitzgerald S;Gil L;Girón CG;Gordon L;Hourlier T;Hunt SE;Janacek SH;Johnson N;Juettemann T;Kähäri AK;Keenan S;Martin FJ;Maurel T;McLaren W;Murphy DN;Nag R;Overduin B;Parker A;Patricio M;Perry E;Pignatelli M;Riat HS;Sheppard D;Taylor K;Thormann A;Vullo A;Wilder SP;Zadissa A;Aken BL;Birney E;Harrow J;Kinsella R;Muffato M;Ruffier M;Searle SM;Spudich G;Trevanion SJ;Yates A;Zerbino DR;Flicek P
通讯作者:
Flicek P
影响因子:
14.9
作者:
Wang K;Li M;Hakonarson H
通讯作者:
Hakonarson H