Predicting runtimes of bioinformatics tools based on historical data: five years of Galaxy usage.

Predicting runtimes of bioinformatics tools based on historical data: five years of Galaxy usage.
复制标题

根据历史数据预测生物信息学工具的运行时间:五年的 Galaxy 使用情况。

DOI:
10.1093/bioinformatics/btz054
复制
发表时间:
2019
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Nekrutenko,Anton
Nekrutenko,Anton
中科院分区:
--
文献类型:
--
作者:
Tyryshkina,Anastasia;Coraor,Nate;Nekrutenko,Anton

文献摘要

参考文献

相似文献

动机在大规模安排生物信息学分析时出现的众多技术挑战之一是确定适当的内存和处理资源数量。分配过多和分配不足都会导致计算基础设施的低效使用。过度分配锁定了原本可以用于其他分析的资源。分配不足会导致作业失败,并需要使用更大的内存或运行时间余量重复分析。我们通过使用Galaxy平台上运行的生物信息学分析的历史数据集来演示在线资源需求评估服务的可行性。结果在这里,我们介绍了Galaxy作业运行数据集,并在资源使用预测任务中测试了流行的机器学习模型。我们包括了三种流行的森林模型:额外树木回归模型、梯度提升回归模型和随机森林回归模型,并发现随机森林在运行时预测任务中表现最好。我们还提供了两种为以前未见过的工作选择工作时间的方法。分位数回归森林的预测更准确,并能够通过改变估计的置信度来提高性能。然而,置信度区间的大小是可变的,不能绝对约束。随机森林分类器通过提供对预测区间大小的控制来解决这一问题,其精度可与回归因子的精度相媲美。我们展示了使用相同的方法来估计作业的内存需求是可能的,就我们所知,这是以前从未做过的。这样的估计对准确的资源分配非常有益。可用性和实现在https://github.com/atyryshkina/algorithm-performance-analysis,上提供了用Python实现的源代码。补充信息补充数据可在BioInformation在线上获得。
MotivationOne of the many technical challenges that arises when scheduling bioinformatics analyses at scale is determining the appropriate amount of memory and processing resources. Both over- and under-allocation leads to an inefficient use of computational infrastructure. Over allocation locks resources that could otherwise be used for other analyses. Under-allocation causes job failure and requires analyses to be repeated with a larger memory or runtime allowance. We address this challenge by using a historical dataset of bioinformatics analyses run on the Galaxy platform to demonstrate the feasibility of an online service for resource requirement estimation.ResultsHere we introduced the Galaxy job run dataset and tested popular machine learning models on the task of resource usage prediction. We include three popular forest models: the extra trees regressor, the gradient boosting regressor and the random forest regressor, and find that random forests perform best in the runtime prediction task. We also present two methods of choosing walltimes for previously unseen jobs. Quantile regression forests are more accurate in their predictions, and grant the ability to improve performance by changing the confidence of the estimates. However, the sizes of the confidence intervals are variable and cannot be absolutely constrained. Random forest classifiers address this problem by providing control over the size of the prediction intervals with an accuracy that is comparable to that of the regressor. We show that estimating the memory requirements of a job is possible using the same methods, which as far as we know, has not been done before. Such estimation can be highly beneficial for accurate resource allocation.Availability and implementationSource code available at https://github.com/atyryshkina/algorithm-performance-analysis, implemented in Python.Supplementary informationSupplementary data are available atBioinformaticsonline.
长期佛波酯治疗可将磷脂酶 D 的激活与缓激肽刺激的内皮细胞中的磷酸肌醇水解和前列环素合成分离。
DOI: 10.1016/0006-291x(89)91072-3
发表时间: 1989
影响因子: 3.1
作者:
Martin,TW;Feldman,DR;Goldstein,KE;Wagner,JR
通讯作者: Wagner,JR
对照和醛固酮高血压大鼠胸主动脉中 α-1 肾上腺素能受体的特征:放射性配体结合与钾流出和收缩的相关性。
DOI: --
发表时间: 1987
期刊: The Journal of pharmacology and experimental therapeutics
影响因子: --
作者:
Smith,JM;Jones,SB;Bylund,DB;Jones,AW
通讯作者: Jones,AW
DOI: --
发表时间: 1982
期刊: Prostaglandins
影响因子: --
作者:
E. Pipili;N. Poyser
通讯作者: N. Poyser
血小板中肌醇磷脂周转率的测量。
DOI: --
发表时间: 1987
影响因子: --
作者:
E. Lapetina;W. Siess
通讯作者: W. Siess
通过内皮细胞中磷脂酰胆碱特异性的磷脂酶 D-磷脂酸磷酸酶途径形成二酰甘油。
DOI: --
发表时间: 1988
期刊: Biochimica et Biophysica Acta
影响因子: --
作者:
T. W. Martin
通讯作者: T. W. Martin