Octopus-toolkit: a workflow to automate mining of public epigenomic and transcriptomic next-generation sequencing data.

Octopus-toolkit: a workflow to automate mining of public epigenomic and transcriptomic next-generation sequencing data.
复制标题

DOI:
10.1093/nar/gky083
复制
发表时间:
2018-05-18
影响因子:
14.9
通讯作者:
Kang K
Kang K
中科院分区:
生物学2区
文献类型:
--
作者:
Kim T;Seo HD;Hennighausen L;Lee D;Kang K

文献摘要

参考文献

被引文献

相似文献

Octopus-TOOLKIT是一个独立的应用程序,只需一个步骤就可以检索和处理大量的下一代测序(NGS)数据。Octopus-TOOLKIT是一个利用Aspera、SRA工具包、FastQC、Trimomatic、HISAT2、STAR、SamTools和HOMER应用程序的自动化设置和分析管道。当程序启动时,所有应用程序都安装在用户的计算机上。在安装后,它可以自动从基因表达总集数据库中检索各种表观基因组和转录组数据集的原始文件,包括ChIP-seq、atac-seq、DNase-seq、MeDIP-seq、MNase-seq和RNA-seq。然后可以按顺序处理下载的文件以生成BAM和Bigwig文件,这些文件用于高级分析和可视化。目前,它可以处理常见的模式基因组,如人(智人)、小鼠(小鼠)、狗(犬狼疮)、植物(拟南芥)、斑马鱼(Danio Rerio)、果蝇(黑腹果蝇)、蠕虫(线虫)和发芽酵母(酿酒酵母)的基因组。利用Octopus-TOOLKIT的处理文件,用户只需很少的命令就可以轻松地进行各种数据集的荟萃分析、DNA结合蛋白的Motif搜索以及差异表达基因和/或蛋白质结合位点的鉴定。总体而言,Octopus-TOOLKIT促进了对现有表观基因组和转录组大数据的系统和综合分析。
Octopus-toolkit is a stand-alone application for retrieving and processing large sets of next-generation sequencing (NGS) data with a single step. Octopus-toolkit is an automated set-up-and-analysis pipeline utilizing the Aspera, SRA Toolkit, FastQC, Trimmomatic, HISAT2, STAR, Samtools, and HOMER applications. All the applications are installed on the user's computer when the program starts. Upon the installation, it can automatically retrieve original files of various epigenomic and transcriptomic data sets, including ChIP-seq, ATAC-seq, DNase-seq, MeDIP-seq, MNase-seq and RNA-seq, from the gene expression omnibus data repository. The downloaded files can then be sequentially processed to generate BAM and BigWig files, which are used for advanced analyses and visualization. Currently, it can process NGS data from popular model genomes such as, human (Homo sapiens), mouse (Mus musculus), dog (Canis lupus familiaris), plant (Arabidopsis thaliana), zebrafish (Danio rerio), fruit fly (Drosophila melanogaster), worm (Caenorhabditis elegans), and budding yeast (Saccharomyces cerevisiae) genomes. With the processed files from Octopus-toolkit, the meta-analysis of various data sets, motif searches for DNA-binding proteins, and the identification of differentially expressed genes and/or protein-binding sites can be easily conducted with few commands by users. Overall, Octopus-toolkit facilitates the systematic and integrative analysis of available epigenomic and transcriptomic NGS big data.
DOI: 10.1016/j.cell.2013.09.053
发表时间: 2013-11-07
期刊: Cell
影响因子: 64.5
作者:
Hnisz D;Abraham BJ;Lee TI;Lau A;Saint-André V;Sigova AA;Hoke HA;Young RA
通讯作者: Young RA
DOI: 10.1016/j.molcel.2010.05.004
发表时间: 2010-05-28
期刊: Molecular cell
影响因子: 16
作者:
Heinz S;Benner C;Spann N;Bertolino E;Lin YC;Laslo P;Cheng JX;Murre C;Singh H;Glass CK
通讯作者: Glass CK
DOI: 10.1534/g3.113.010140
发表时间: 2014-04-16
期刊: G3 (Bethesda, Md.)
影响因子: --
作者:
Salas-Santiago B;Lopes JM
通讯作者: Lopes JM
DOI: 10.1101/gr.4086505
发表时间: 2005-10-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Giardine, B;Riemer, C;Nekrutenko, A
通讯作者: Nekrutenko, A
DOI: 10.1016/j.cell.2008.02.022
发表时间: 2008-03-07
期刊: CELL
影响因子: 64.5
作者:
Schones, Dustin E.;Cui, Kairong;Zhao, Keji
通讯作者: Zhao, Keji