GAUGE-Annotated Microbial Transcriptomic Data Facilitate Parallel Mining and High-Throughput Reanalysis To Form Data-Driven Hypotheses.

GAUGE-Annotated Microbial Transcriptomic Data Facilitate Parallel Mining and High-Throughput Reanalysis To Form Data-Driven Hypotheses.
复制标题

规范注释的微生物转录组数据促进并行挖掘和高重复性再分析,以形成数据驱动的假设。

DOI:
10.1128/msystems.01305-20
复制
发表时间:
2021-03-23
期刊:
影响因子:
6.4
通讯作者:
Hampton TH
Hampton TH
中科院分区:
生物学2区
文献类型:
--
作者:
Li Z;Koeppen K;Holden VI;Neff SL;Cengher L;Demers EG;Mould DL;Stanton BA;Hampton TH

文献摘要

被引文献

相似文献

NCBI Gene Expression Omnibus(GEO)提供了查询和下载转录组数据的工具。然而,只有不到4%的微生物实验包括评估差异基因表达以进行高通量再分析所需的样本组注释,并且2014年之后保存的数据普遍缺乏这些注释。我们的算法GAUGE(使用文本/数据组集成的一般注释)自动注释GEO微生物数据集,包括微阵列和RNA测序研究,将适合分析的数据集的百分比从4%增加到33%。89%的GAUGE注释研究与人类策展人生成的小组任务相匹配。为了证明GAUGE注释如何带来科学见解,我们创建了GAPE(GAUGE注释的铜绿假单胞菌和大肠杆菌转录组学再分析纲要),这是一个Shiny Web界面,用于分析73个GAUGE注释的铜绿假单胞菌研究,比以前多三倍。GAPE分析显示,PA3923,一个未知功能的基因,经常在超过50%的研究中差异表达,并与生物膜形成相关的基因显着共调节。后续的湿台实验表明,PA3923突变体在生物膜形成方面确实存在缺陷,这与GAUGE和GAPE促进的预测一致。我们预计,我们免费提供的GAUGE和GAPE将使公开的微生物转录组数据更容易重复使用,并导致新的数据驱动的假设。GEO存档了来自5,800多个微生物实验的转录组数据,并允许研究人员回答发表论文中没有直接解决的问题。然而,不到4%的微生物数据集包括高通量再分析所需的样品组注释。这种限制阻碍了大量的微生物转录组数据被容易地重复使用。在这里,我们证明了GAUGE算法可以使33%的微生物数据可用于并行挖掘和重新分析。GAUGE注释增加了统计能力,从而使差异基因表达的一致模式更容易识别。此外,我们还开发了GAPE(GAUGE注释的铜绿假单胞菌和大肠杆菌转录组再分析纲要),这是一个对铜绿假单胞菌和大肠杆菌进行并行分析的Shiny Web界面。大肠杆菌药典GAUGE和GAPE的源代码是免费提供的,可以重新用于创建其他细菌物种的纲要。作者视频:本文的作者视频摘要可用。
The NCBI Gene Expression Omnibus (GEO) provides tools to query and download transcriptomic data. However, less than 4% of microbial experiments include the sample group annotations required to assess differential gene expression for high-throughput reanalysis, and data deposited after 2014 universally lack these annotations. Our algorithm GAUGE (general annotation using text/data group ensembles) automatically annotates GEO microbial data sets, including microarray and RNA sequencing studies, increasing the percentage of data sets amenable to analysis from 4% to 33%. Eighty-nine percent of GAUGE-annotated studies matched group assignments generated by human curators. To demonstrate how GAUGE annotation can lead to scientific insight, we created GAPE (GAUGE-annotated Pseudomonas aeruginosa and Escherichia coli transcriptomic compendia for reanalysis), a Shiny Web interface to analyze 73 GAUGE-annotated P. aeruginosa studies, three times more than previously available. GAPE analysis revealed that PA3923, a gene of unknown function, was frequently differentially expressed in more than 50% of studies and significantly coregulated with genes involved in biofilm formation. Follow-up wet-bench experiments demonstrate that PA3923 mutants are indeed defective in biofilm formation, consistent with predictions facilitated by GAUGE and GAPE. We anticipate that GAUGE and GAPE, which we have made freely available, will make publicly available microbial transcriptomic data easier to reuse and lead to new data-driven hypotheses. IMPORTANCE GEO archives transcriptomic data from over 5,800 microbial experiments and allows researchers to answer questions not directly addressed in published papers. However, less than 4% of the microbial data sets include the sample group annotations required for high-throughput reanalysis. This limitation blocks a considerable amount of microbial transcriptomic data from being reused easily. Here, we demonstrate that the GAUGE algorithm could make 33% of microbial data accessible to parallel mining and reanalysis. GAUGE annotations increase statistical power and, thereby, make consistent patterns of differential gene expression easier to identify. In addition, we developed GAPE (GAUGE-annotated Pseudomonas aeruginosa and Escherichia coli transcriptomic compendia for reanalysis), a Shiny Web interface that performs parallel analyses on P. aeruginosa and E. coli compendia. Source code for GAUGE and GAPE is freely available and can be repurposed to create compendia for other bacterial species. Author Video: An author video summary of this article is available.