CBS Genome Atlas Database: a dynamic storage for bioinformatic results and sequence data

CBS Genome Atlas Database: a dynamic storage for bioinformatic results and sequence data
复制标题

DOI:
10.1093/bioinformatics/bth423
复制
发表时间:
2004-12-12
期刊:
影响因子:
5.8
通讯作者:
Ussery, DW
Ussery, DW
中科院分区:
生物学3区
文献类型:
--
作者:
Hallin, PF;Ussery, DW

文献摘要

被引文献

相似文献

目前,新的细菌基因组正在每月公布。随着基因组序列数据量的增长,需要一种灵活且易于维护的结构来存储序列数据和生物信息学分析的结果。现在已经有150多个测序的细菌基因组,许多生物学家还不能很容易地比较分类学上相似的生物体的特性。除了AT含量、染色体长度、tRNA计数和rRNA计数等最基本的信息外,还需要进行大量更复杂的计算才能进行详细的比较基因组学研究。DNA结构计算,如曲率和堆积能,DNA组成,如碱基偏斜,寡核苷酸偏斜和重复在本地和全球水平上只是一些分析,在哥伦比亚广播公司基因组图谱网页上提出。复杂的分析、不断变化的方法和频繁添加新模型是需要动态数据库布局的因素。使用GNU Make系统、csh、Perl和MySQL等基本工具,我们创建了一个灵活的数据库环境,用于存储和维护完整微生物基因组集合的结果。目前,这些结果统计到220多条信息。该解决方案的主干由一个用Perl编写的程序包组成,它使管理员能够同步和更新数据库内容。MySQL数据库已通过PHP4连接到CBS Web服务器,为中心外的用户提供动态Web内容。这个解决方案与现有的服务器基础设施紧密配合,这里提出的解决方案可能可以作为其他研究小组解决数据库问题的模板。
Currently, new bacterial genomes are being published on a monthly basis. With the growing amount of genome sequence data, there is a demand for a flexible and easy-to-maintain structure for storing sequence data and results from bioinformatic analysis. More than 150 sequenced bacterial genomes are now available, and comparisons of properties for taxonomically similar organisms are not readily available to many biologists. In addition to the most basic information, such as AT content, chromosome length, tRNA count and rRNA count, a large number of more complex calculations are needed to perform detailed comparative genomics. DNA structural calculations like curvature and stacking energy, DNA compositions like base skews, oligo skews and repeats at the local and global level are just a few of the analysis that are presented on the CBS Genome Atlas Web page. Complex analysis, changing methods and frequent addition of new models are factors that require a dynamic database layout. Using basic tools like the GNU Make system, csh, Perl and MySQL, we have created a flexible database environment for storing and maintaining such results for a collection of complete microbial genomes. Currently, these results counts to more than 220 pieces of information. The backbone of this solution consists of a program package written in Perl, which enables administrators to synchronize and update the database content. The MySQL database has been connected to the CBS web-server via PHP4, to present a dynamic web content for users outside the center. This solution is tightly fitted to existing server infrastructure and the solutions proposed here can perhaps serve as a template for other research groups to solve database issues.