CMS Analysis and Data Reduction with Apache Spark

CMS Analysis and Data Reduction with Apache Spark
复制标题

使用 Apache Spark 进行 CMS 分析和数据缩减

DOI:
--
复制
发表时间:
2017
期刊:
Journal of Physics: Conference Series
影响因子:
--
通讯作者:
Alexey Svyatkovskiy
Alexey Svyatkovskiy
中科院分区:
--
文献类型:
--
作者:
O. Gutsche;L. Canali;Illia Cremer;M. Cremonesi;P. Elmer;I. Fisk;M. Girone;B. Jayatilaka;J. Kowalkowski;V. Khristenko;Evangelos Motesnitsalis;J. Pivarski;S. Sehrish;K. Surdy;Alexey Svyatkovskiy

文献摘要

被引文献

相似文献

几十年来,实验粒子物理学一直处于分析世界上最大数据集的最前沿。HEP社区是最早为这项任务开发合适的软件和计算工具的社区之一。近年来,用于分布式数据处理的新工具包和系统(统称为“大数据”技术)已经从工业和开源项目中出现,以支持工业中的PB和EB数据集的分析。虽然HEP中的数据分析原则没有改变(过滤和转换实验特定的数据格式),但这些新技术使用不同的方法和工具,有望以全新的眼光分析非常大的数据集,从而可能减少物理时间,增加交互性。此外,这些新工具通常由大型社区积极开发,通常利用行业资源,并在开源许可下开发。这些因素促进了工具的采用和成熟,并促进了支持它们的社区,同时有助于降低最终用户的拥有成本。在这次演讲中,我们将介绍使用Apache Spark进行最终用户数据分析的研究。我们正在研究HEP分析工作流程,分为两个重点:减少集中产生的实验数据集和最终分析到出版图。研究第一个推力,CMS正在与CERN openlab和英特尔合作开发CMS大数据缩减设施。目标是将1 PB的官方CMS数据减少到1 TB的ntuple输出以供分析。我们正在介绍这个为期2年的项目的进展情况,以及基于Spark的HEP分析的初步结果。研究第二个推力,我们正在研究使用Apache Spark进行CMS暗物质物理搜索,与传统的基于根的分析相比,研究Spark的可行性,可用性和性能。
Experimental Particle Physics has been at the forefront of analyzing the world’s largest datasets for decades. The HEP community was among the first to develop suitable software and computing tools for this task. In recent times, new toolkits and systems for distributed data processing, collectively called “Big Data” technologies have emerged from industry and open source projects to support the analysis of Petabyte and Exabyte datasets in industry. While the principles of data analysis in HEP have not changed (filtering and transforming experiment-specific data formats), these new technologies use different approaches and tools, promising a fresh look at analysis of very large datasets that could potentially reduce the time-to-physics with increased interactivity. Moreover these new tools are typically actively developed by large communities, often profiting of industry resources, and under open source licensing. These factors result in a boost for adoption and maturity of the tools and for the communities supporting them, at the same time helping in reducing the cost of ownership for the end users. In this talk, we are presenting studies of using Apache Spark for end user data analysis. We are studying the HEP analysis workflow separated into two thrusts: the reduction of centrally produced experiment datasets and the end analysis up to the publication plot. Studying the first thrust, CMS is working together with CERN openlab and Intel on the CMS Big Data Reduction Facility. The goal is to reduce 1 PB of official CMS data to 1 TB of ntuple output for analysis. We are presenting the progress of this 2-year project with first results of scaling up Spark-based HEP analysis. Studying the second thrust, we are presenting studies on using Apache Spark for a CMS Dark Matter physics search, investigating Spark’s feasibility, usability and performance compared to the traditional ROOT-based analysis.