ParPEST: a pipeline for EST data analysis based on parallel computing

ParPEST: a pipeline for EST data analysis based on parallel computing
复制标题

DOI:
10.1186/1471-2105-6-s4-s9
复制
发表时间:
2005-12-01
期刊:
影响因子:
3
通讯作者:
Chiusano, ML
Chiusano, ML
中科院分区:
生物学4区
文献类型:
--
作者:
D'Agostino, N;Aversano, M;Chiusano, ML

文献摘要

被引文献

相似文献

背景:表达序列标签(expressedsequencetags,ESTs)是从随机选择的cDNA克隆的5'端和3'端产生的短且易出错的DNA序列。它们为比较和功能基因组学研究提供了重要的资源,并且为基因组序列的注释提供了可靠的信息。由于生物技术的进步,EST每天都以大型数据集的形式确定。因此,合适的和有效的生物信息学方法是必要的,以组织数据相关的信息内容为进一步investigation.Results:我们实现了ParPEST(并行处理的EST),一个管道的基础上并行计算EST分析。将结果组织在合适的数据仓库中,以提供挖掘表达序列数据集的起点。收集的信息是有用的调查数据质量和数据信息的内容,也丰富了初步的功能annotation.Conclusion:这里提出的管道已经开发到执行一个详尽的和可靠的分析EST数据,并提供一个精心策划的一套基于关系数据库的信息。此外,它旨在减少使用分布式流程和并行软件进行完整分析所需的特定步骤的执行时间。它被设想为在低要求的硬件组件上运行,以满足日益增长的需求,典型的数据使用,以及可扩展性,成本低廉。
Background: Expressed Sequence Tags (ESTs) are short and error-prone DNA sequences generated from the 5' and 3' ends of randomly selected cDNA clones. They provide an important resource for comparative and functional genomic studies and, moreover, represent a reliable information for the annotation of genomic sequences. Because of the advances in biotechnologies, ESTs are daily determined in the form of large datasets. Therefore, suitable and efficient bioinformatic approaches are necessary to organize data related information content for further investigations.Results: We implemented ParPEST (Parallel Processing of ESTs), a pipeline based on parallel computing for EST analysis. The results are organized in a suitable data warehouse to provide a starting point to mine expressed sequence datasets. The collected information is useful for investigations on data quality and on data information content, enriched also by a preliminary functional annotation.Conclusion: The pipeline presented here has been developed to perform an exhaustive and reliable analysis on EST data and to provide a curated set of information based on a relational database. Moreover, it is designed to reduce execution time of the specific steps required for a complete analysis using distributed processes and parallelized software. It is conceived to run on low requiring hardware components, to fulfill increasing demand, typical of the data used, and scalability at affordable costs.