DFAST and DAGA: web-based integrated genome annotation tools and resources.

DFAST and DAGA: web-based integrated genome annotation tools and resources.
复制标题

DFAST和DAGA:基于Web的集成基因组注释工具和资源。

DOI:
10.12938/bmfh.16-003
复制
发表时间:
2016
期刊:
Bioscience of microbiota, food and health
影响因子:
--
通讯作者:
Arita M
Arita M
中科院分区:
其他
文献类型:
--
作者:
Tanizawa Y;Fujisawa T;Kaminuma E;Nakamura Y;Arita M

文献摘要

参考文献

被引文献

相似文献

公共序列数据库的数据质量保证和正确的分类归属一直是一个永恒的问题。DDBJ快速注释和提交工具(DFAST)是一个新开发的基因组注释管道,具有质量和分类评估工具。为了能够注释即可提交的质量,我们还构建了为乳酸菌量身定制的精选参考蛋白质数据库。DFAST的开发是为了使DDBJ提交所需的所有程序都可以在线无缝完成。在线工作空间对于不熟悉生物信息学技能的用户特别有用。此外,我们还开发了一个基因组库,DFAST Archive of Genome Annotation(DAGA),目前包括1,421个基因组,涵盖乳杆菌属和片球菌属两个属的179个物种和18个亚种,从DDBJ/ENA/GenBank和序列读取档案(SRA)获得。保存在DAGA中的所有基因组都进行了一致的注释,并使用DFAST进行评估。为了评估基于基因组序列信息的分类位置,我们使用了平均核苷酸同一性(ANI),它显示出很高的区分能力,以确定两个给定的基因组是否属于同一物种。我们纠正了公共数据库中错误标记或错误识别的基因组,并将精选的信息保存在DAGA中。该库将提高乳酸菌基因组资源的可访问性和可重复使用性。通过利用DAGA中保存的数据,我们发现了加氏乳杆菌和詹氏乳杆菌的种内亚群,其亚群之间的变化大于公认的ANI阈值95%,以区分物种。DFAST和DAGA可在https://dfast.nig.ac.jp上免费访问。
Quality assurance and correct taxonomic affiliation of data submitted to public sequence databases have been an everlasting problem. The DDBJ Fast Annotation and Submission Tool (DFAST) is a newly developed genome annotation pipeline with quality and taxonomy assessment tools. To enable annotation of ready-to-submit quality, we also constructed curated reference protein databases tailored for lactic acid bacteria. DFAST was developed so that all the procedures required for DDBJ submission could be done seamlessly online. The online workspace would be especially useful for users not familiar with bioinformatics skills. In addition, we have developed a genome repository, DFAST Archive of Genome Annotation (DAGA), which currently includes 1,421 genomes covering 179 species and 18 subspecies of two genera, Lactobacillus and Pediococcus, obtained from both DDBJ/ENA/GenBank and Sequence Read Archive (SRA). All the genomes deposited in DAGA were annotated consistently and assessed using DFAST. To assess the taxonomic position based on genomic sequence information, we used the average nucleotide identity (ANI), which showed high discriminative power to determine whether two given genomes belong to the same species. We corrected mislabeled or misidentified genomes in the public database and deposited the curated information in DAGA. The repository will improve the accessibility and reusability of genome resources for lactic acid bacteria. By exploiting the data deposited in DAGA, we found intraspecific subgroups in Lactobacillus gasseri and Lactobacillus jensenii, whose variation between subgroups is larger than the well-accepted ANI threshold of 95% to differentiate species. DFAST and DAGA are freely accessible at https://dfast.nig.ac.jp.
DOI: 10.1093/nar/gkv1323
发表时间: 2016-01-04
影响因子: 14.9
作者:
Cochrane G;Karsch-Mizrachi I;Takagi T;International Nucleotide Sequence Database Collaboration
通讯作者: International Nucleotide Sequence Database Collaboration
DOI: 10.1093/nar/gkr854
发表时间: 2012-01
影响因子: 14.9
作者:
Kodama Y;Shumway M;Leinonen R;International Nucleotide Sequence Database Collaboration
通讯作者: International Nucleotide Sequence Database Collaboration
DOI: 10.1186/s40793-016-0134-1
发表时间: 2016-02-09
影响因子: --
作者:
Federhen S;Rossello-Mora R;Klenk HP;Tindall BJ;Konstantinidis KT;Whitman WB;Brown D;Labeda D;Ussery D;Garrity GM;Colwell RR;Hasan N;Graf J;Parte A;Yarza P;Goldberg B;Sichtig H;Karsch-Mizrachi I;Clark K;McVeigh R;Pruitt KD;Tatusova T;Falk R;Turner S;Madden T;Kitts P;Kimchi A;Klimke W;Agarwala R;DiCuccio M;Ostell J
通讯作者: Ostell J
DOI: 10.1099/ijs.0.059980-0
发表时间: 2014-09-01
影响因子: 2.8
作者:
Isabel Puertas, Ana;Arahal, David R.;Teresa Duenas, M.
通讯作者: Teresa Duenas, M.
DOI: 10.1371/journal.pone.0077910
发表时间: 2013
期刊: PloS one
影响因子: 3.7
作者:
Nakazato T;Ohta T;Bono H
通讯作者: Bono H