EST2uni: an open, parallel tool for automated EST analysis and database creation, with a data mining web interface and microarray expression data integration.

EST2uni: an open, parallel tool for automated EST analysis and database creation, with a data mining web interface and microarray expression data integration.
复制标题

DOI:
10.1186/1471-2105-9-5
复制
发表时间:
2008-01-07
期刊:
影响因子:
3
通讯作者:
Blanca JM
Blanca JM
中科院分区:
生物学4区
文献类型:
--
作者:
Forment J;Gilabert F;Robles A;Conejero V;Nuez F;Blanca JM

文献摘要

参考文献

被引文献

相似文献

表达序列标签(EST)集合由大量的单遍、冗余的部分序列组成,需要对这些序列进行处理、聚类和注释,以去除低质量和载体区域,消除冗余和测序错误,并提供生物学相关信息。为了提供一种适当的方法来执行分析无害环境技术的不同步骤,必须发展适应具体无害环境技术项目当地需要的灵活的计算管道。此外,EST集合必须存储在高度结构化的关系数据库中,研究人员可以通过用户友好的界面获得这些数据库,这些界面允许有效和复杂的数据挖掘,从而为它们的充分利用提供最大的能力。我们已经创建了EST2uni,这是一个集成的、高度可配置的EST分析管道和数据挖掘软件包,可以自动处理EST集合的预处理、聚类、注释、数据库创建和数据挖掘。该管道使用标准的EST分析工具,软件采用模块化设计,便于添加新的分析方法及其配置。目前实施的分析包括功能和结构注释,SNP和微卫星发现,先前已知遗传标记数据和基因表达结果的整合,以及cDNA微阵列设计的协助。它可以在PC集群中并行运行,以减少分析所需的时间。它还创建了一个链接到数据库的网站,显示收集的统计数据,具有复杂的查询功能和数据挖掘和检索工具。该软件包提供了一个高效、完整的生物信息学管理工具,可以很容易地适应不同EST项目的本地需求。代码在GPL许可下是免费的,可以在。本网站还提供了安装和配置软件包的详细说明。该代码正在积极开发中,以纳入新的分析,方法和算法,因为它们是由生物信息学社区发布的。
Expressed sequence tag (EST) collections are composed of a high number of single-pass, redundant, partial sequences, which need to be processed, clustered, and annotated to remove low-quality and vector regions, eliminate redundancy and sequencing errors, and provide biologically relevant information. In order to provide a suitable way of performing the different steps in the analysis of the ESTs, flexible computation pipelines adapted to the local needs of specific EST projects have to be developed. Furthermore, EST collections must be stored in highly structured relational databases available to researchers through user-friendly interfaces which allow efficient and complex data mining, thus offering maximum capabilities for their full exploitation. We have created EST2uni, an integrated, highly-configurable EST analysis pipeline and data mining software package that automates the pre-processing, clustering, annotation, database creation, and data mining of EST collections. The pipeline uses standard EST analysis tools and the software has a modular design to facilitate the addition of new analytical methods and their configuration. Currently implemented analyses include functional and structural annotation, SNP and microsatellite discovery, integration of previously known genetic marker data and gene expression results, and assistance in cDNA microarray design. It can be run in parallel in a PC cluster in order to reduce the time necessary for the analysis. It also creates a web site linked to the database, showing collection statistics, with complex query capabilities and tools for data mining and retrieval. The software package presented here provides an efficient and complete bioinformatics tool for the management of EST collections which is very easy to adapt to the local needs of different EST projects. The code is freely available under the GPL license and can be obtained at . This site also provides detailed instructions for installation and configuration of the software package. The code is under active development to incorporate new analyses, methods, and algorithms as they are released by the bioinformatics community.
DOI: 10.1186/1471-2105-6-31
发表时间: 2005-02-15
期刊: BMC bioinformatics
影响因子: 3
作者:
Slater GS;Birney E
通讯作者: Birney E
DOI: 10.2144/01315dd03
发表时间: 2001-11-01
期刊: BIOTECHNIQUES
影响因子: 2.7
作者:
Evertsz, EM;Au-Young, J;Reynolds, MA
通讯作者: Reynolds, MA
DOI: 10.1038/ng0893-373
发表时间: 1993-08-01
期刊: NATURE GENETICS
影响因子: 30.8
作者:
ADAMS, MD;SOARES, MB;VENTER, JC
通讯作者: VENTER, JC
DOI: 10.1093/bioinformatics/btg205
发表时间: 2003-09-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Mao, CH;Cushman, JC;Weller, JW
通讯作者: Weller, JW
DOI: 10.1186/1471-2105-6-s4-s9
发表时间: 2005-12-01
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
D'Agostino, N;Aversano, M;Chiusano, ML
通讯作者: Chiusano, ML