Framework: Awkward Arrays - Accelerating scientific data analysis on irregularly shaped data
Framework: Awkward Arrays - Accelerating scientific data analysis on irregularly shaped data
批准号:
2103945
负责人:
Jim Pivarski
金额:
$106.75万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2021
资助国家:
美国
项目状态:
已结题
起止时间:
2021-09-01 至 2024-08-31
中文摘要
如今,科学家不仅要成为各自领域的专家,还要成为程序员。任何简化分析数据的计算体操的软件都是受欢迎的,因为它允许科学家将更多的注意力集中在他们想要计算的东西上,而不是如何计算,但是在易用性和计算速度之间通常存在紧张关系。更快的计算意味着分析更多的数据,但最重要的是没有错误的代码。Awkward Array是一个软件库,它可以对不适合整齐行的数据进行计算。例如,一组海底探测器可能在不同深度测量海洋温度,但像Excel、Pandas和SQL这样的数据分析工具需要将测量结果排列在矩形表中,每行有相同数量的列。当变量数量的度量嵌套在变量数量的实体中时,问题会更加严重。传统上,科学家们要么使用缓慢的脚本语言,要么使用无情的“裸机”语言来对这些非表格数据集进行计算。Awkward Array对数组概念进行了一般化,使得曾经仅应用于矩形表的简单快速表达式现在可以应用于不规则形状的数据,从而在高级脚本语言中实现了极快的速度。这个项目扩大了笨拙阵列的适用性,超出了它最初打算用于粒子物理学的问题领域,扩展到各种科学领域,包括海洋学、天文学、遗传学、化学和卫生保健。它还将Awkward Array库与流行的数据科学和机器学习工具集成在一起,并将实现扩展到gpu,以实现对相同数组习惯的极快处理。科学家们使用基于numpy的工具,如Pandas,来有效地分析大型和有规律形状的数据。将数字表打包到连续数组中允许操作被预编译和快速;将操作表示为带有隐式循环的简明命令,使它们在数据探索过程中更容易阅读和更快地键入。然而,科学数据往往具有复杂性,无法很好地适应表格格式。这迫使科学家编写带有显式循环的程序,这在Python中很慢。但是,如何分析大型和不规则形状的数据呢?Awkward Array是NumPy的泛化;它是一个Python库,定义了具有任意类型的对象数组以及对其进行的一套泛型操作。像NumPy数组一样,笨拙数组被打包到连续的缓冲区中以提高效率,但与NumPy数组不同的是,它们可以包括变长列表、嵌套字段、缺失值和混合类型。笨拙数组使用的内存比同等Python对象少10倍,计算速度快100倍。该项目将笨拙阵列推广为科学基础图书馆。这是一项跨学科的努力,扩展了Awkward Array,以有效地解决科学家在各种领域和工业数据科学中面临的挑战。项目成员与科学合作者一起解决特定问题,必要时为Awkward Array添加功能,Anaconda和Nvidia的行业合作者更新Python的标准工具,以识别CPU和GPU工作负载的这些新数组类型。它也是一个教育项目,向实践中的科学家和学生教授使用数组习语解决复杂问题的新方法。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Nowadays, scientists need to be programmers as well as experts in their fields. Any software that simplifies the computational gymnastics of analyzing data is welcomed, as it allows scientists to focus more of their attention on what they want to compute, rather than how, but there is usually a tension between ease of use and computational speed. Faster computation means more data analyzed, but error-free code matters most. Awkward Array is a software library that performs calculations on data that do not fit into neat rows. For instance, a fleet of undersea probes may each measure ocean temperatures at a different number of depths, but data analysis tools like Excel, Pandas, and SQL require measurements to be arranged in rectangular tables with the same number of columns in each row. The problem is more acute when variable numbers of measurements are nested within variable numbers of entities. Traditionally, scientists have either used slow scripting languages or unforgiving "bare metal" languages to perform calculations on these non-tabular datasets. Awkward Array generalizes array concepts so that the easy and fast expressions that once applied only to rectangular tables now apply to irregularly shaped data, allowing bare metal speed in a high-level scripting language. This project broadens the applicability of Awkward Array beyond the problem domain it was originally intended for, particle physics, to a wide variety of scientific fields, including oceanography, astronomy, genetics, chemistry, and health care. It also integrates the Awkward Array library with popular data science and machine learning tools and extends the implementation to GPUs for extremely fast processing of the same array idioms.Scientists use NumPy-based tools, such as Pandas, to analyze large and regularly shaped data efficiently. Packing tables of numbers into contiguous arrays allows operations to be precompiled and fast; and expressing operations as concise commands with implicit loops makes them easier to read and quicker to type during data exploration. However, scientific data often has a complexity that does not fit well into tabular format. This forces scientists to write programs with explicit loops, which are slow in Python. But what about analysis of large and irregularly shaped data? Awkward Array is a generalization of NumPy; it is a Python library defining arrays of objects with arbitrary types and a suite of generic operations on them. Like NumPy arrays, Awkward Arrays are packed into contiguous buffers for efficiency, but unlike NumPy arrays, they can include variable-length lists, nested fields, missing values, and mixed types. Awkward Arrays use 10 times less memory than equivalent Python objects and are up to 100 times faster in computations. This project generalizes Awkward Array as a foundational library for science. It is a cross-disciplinary effort, extending Awkward Array to efficiently solve challenges faced by scientists in a variety of fields and data science in industry. The project members work with scientific collaborators to solve specific problems, adding features to Awkward Array if necessary, and industry collaborators at Anaconda and Nvidia update Python’s standard tools to recognize these new array types for CPU and GPU workloads. It is also an educational project, teaching new approaches to complex problems using array idioms, both to practicing scientists and to students.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金