The Sleipnir library for computational functional genomics

The Sleipnir library for computational functional genomics
复制标题

DOI:
10.1093/bioinformatics/btn237
复制
发表时间:
2008-07-01
期刊:
影响因子:
5.8
通讯作者:
Troyanskaya, Olga G.
Troyanskaya, Olga G.
中科院分区:
生物学3区
文献类型:
--
作者:
Huttenhower, Curtis;Schroeder, Mark;Troyanskaya, Olga G.

文献摘要

被引文献

相似文献

动机:生物数据的产生已经加速到这样的程度:对于许多模式生物,成百上千种不同类型的全基因组数据集是可用的。当以综合的方式进行分析时,这些丰富的数据可以导致有价值的生物学见解,但是管理如此大的数据集合的计算挑战是实质性的。为了有效地挖掘这些数据,有必要开发使用存储、内存和处理资源的方法。结果:Sleipnir C库实现了各种机器学习和数据操作算法,重点是异构数据集成和非常大的生物数据收集的效率。Sleipnir允许微阵列处理、功能本体挖掘、聚类、贝叶斯学习和推理以及支持向量机任务在以前不实际的规模上执行异构数据。除了可以很容易地集成到新的计算系统中的库之外,还提供了预构建工具来执行各种常见任务。在桌面或高吞吐量计算环境中,许多工具都是多线程并行的,使用标准个人计算机可以在几分钟内执行数百个数据集的大多数任务。
Motivation: Biological data generation has accelerated to the point where hundreds or thousands of whole-genome datasets of various types are available for many model organisms. This wealth of data can lead to valuable biological insights when analyzed in an integrated manner, but the computational challenge of managing such large data collections is substantial. In order to mine these data efficiently, it is necessary to develop methods that use storage, memory and processing resources carefully.Results: The Sleipnir C library implements a variety of machine learning and data manipulation algorithms with a focus on heterogeneous data integration and efficiency for very large biological data collections. Sleipnir allows microarray processing, functional ontology mining, clustering, Bayesian learning and inference and support vector machine tasks to be performed for heterogeneous data on scales not previously practical. In addition to the library, which can easily be integrated into new computational systems, prebuilt tools are provided to perform a variety of common tasks. Many tools are multithreaded for parallelization in desktop or high-throughput computing environments, and most tasks can be performed in minutes for hundreds of datasets using a standard personal computer.