Scientific machine learning benchmarks

Scientific machine learning benchmarks
复制标题

DOI:
10.1038/s42254-022-00441-7
复制
发表时间:
2021-10
影响因子:
38.5
通讯作者:
J. Thiyagalingam;M. Shankar;Geoffrey Fox;Tony (Anthony) John Grenville Hey
J. Thiyagalingam;M. Shankar;Geoffrey Fox;Tony (Anthony) John Grenville Hey
中科院分区:
物理与天体物理1区
文献类型:
--
作者:
J. Thiyagalingam;M. Shankar;Geoffrey Fox;Tony (Anthony) John Grenville Hey

文献摘要

相似文献

深度学习改变了机器学习技术在大型实验数据集分析中的使用。在科学中,此类数据集通常由大规模实验设施生成,而机器学习侧重于识别模式、趋势和异常,以从数据中提取有意义的科学见解。在即将建立的实验设施中,例如英国的极限光子学应用中心(EPAC)或国际平方公里阵列(SKA),数据生成的速率和数据量的规模将越来越需要使用更加自动化的数据分析。然而,目前,由于许多不同的机器学习框架、计算机架构和机器学习模型的潜在适用性,确定最合适的机器学习算法来分析任何给定的科学数据集是一个挑战。从历史上看,对于高性能计算系统的建模和仿真,这些问题是通过对计算机应用程序、算法和架构进行基准测试来解决的。对于科学家和计算机科学家来说,扩展这种基准测试方法并确定将机器学习方法应用于开放的、精心策划的科学数据集的指标是一个新的挑战。在这里,我们介绍科学机器学习基准的概念并回顾现有方法。作为示例,我们描述了科学机器学习基准的 SciMLBench 套件。
Deep learning has transformed the use of machine learning technologies for the analysis of large experimental datasets. In science, such datasets are typically generated by large-scale experimental facilities, and machine learning focuses on the identification of patterns, trends and anomalies to extract meaningful scientific insights from the data. In upcoming experimental facilities, such as the Extreme Photonics Application Centre (EPAC) in the UK or the international Square Kilometre Array (SKA), the rate of data generation and the scale of data volumes will increasingly require the use of more automated data analysis. However, at present, identifying the most appropriate machine learning algorithm for the analysis of any given scientific dataset is a challenge due to the potential applicability of many different machine learning frameworks, computer architectures and machine learning models. Historically, for modelling and simulation on high-performance computing systems, these issues have been addressed through benchmarking computer applications, algorithms and architectures. Extending such a benchmarking approach and identifying metrics for the application of machine learning methods to open, curated scientific datasets is a new challenge for both scientists and computer scientists. Here, we introduce the concept of machine learning benchmarks for science and review existing approaches. As an example, we describe the SciMLBench suite of scientific machine learning benchmarks.