课题基金 / 基金详情

SHF: Medium: Interactive Debegging for Big Data Analytics

SHF: Medium: Interactive Debegging for Big Data Analytics
SHF:中:大数据分析的交互式调试
批准号:
1764077
负责人:
Miryung Kim
金额:
$90.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-07-01 至 2024-06-30

项目摘要

项目成果

Miryung Kim的其他基金

相似基金

相关文献

中文摘要
翻译
科学、工程、国家安全和医疗保健领域的大量数据导致了大数据分析这一新兴领域的出现。为了处理大量数据,开发人员利用云中的数据密集型可伸缩计算(DISC)系统,例如谷歌的MapReduce、Apache Hadoop和Apache Spark。虽然DISC系统有助于解决大数据分析的可扩展性挑战,但它们也为数据科学家在理解和解决错误方面带来了巨大的挑战。这个项目解决了目前DISC系统中严重缺乏调试支持的问题,这使得数据科学家很难理解他们的应用程序,确定识别错误的原因,并确保这些错误得到适当的修复。本研究为Apache Spark等现代DISC系统中的大数据处理程序提供了两种调试支持:用于大规模分布式处理的新型交互式实时调试原语和用于大数据的工具辅助故障定位服务。技术方法包括一种新的数据来源技术,用于为迭代开发和调试工作负载提供对大规模分布式数据处理和运行时优化的细粒度可见性。工具辅助的故障定位服务利用这些潜在的来源和优化技术,有效地查明和描述错误的根本原因。大数据分析在21世纪变得越来越重要,人们的日常生活留下了详细的数字记录,从公司到政府机构的各种决策者都希望以数据为基础采取行动。该研究有助于提高大数据应用的生产力和正确性,这对于许多将tb的低价值数据提炼成高价值见解的学科至关重要。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
An abundance of data in science, engineering, national security, and health care has led to the emerging field of big data analytics. To process massive quantities of data, developers leverage data-intensive scalable computing (DISC) systems in the cloud, such as Google's MapReduce, Apache Hadoop, and Apache Spark. While DISC systems help to address the scalability challenges of big data analytics, they also introduce an enormous challenge for data scientists in understanding and resolving errors. This project addresses the severe lack of debugging support in DISC systems today, which makes it difficult for data scientists to understand their applications, determine the causes of identified errors, and ensure that such errors are properly repaired. The research provides two kinds of debugging support for big data processing programs in modern DISC systems like Apache Spark: new interactive, real-time debugging primitives for large-scale distributed processing and tool-assisted fault-localization services for big data. Technical approaches include a new data provenance technique for providing fine-grained visibility into large-scale distributed data processing and runtime optimizations for iterative development and debugging workloads. Tool-assisted fault localization services leverage these underlying provenance and optimization techniques to pinpoint and characterize the root causes of errors efficiently. Big data analytics is increasingly important in the 21st century, where daily lives leave behind a detailed digital record and decision-makers of all kinds, from companies to government agencies, would like to base their actions on data. The research contributes to improving productivity and correctness of big data applications, which is crucial for many disciplines that distill terabytes of low-value data into high-value insights.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(16)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1109/issre.2019.00020
发表时间: 2019-10
期刊: 2019 IEEE 30th International Symposium on Software Reliability Engineering (ISSRE)
影响因子: --
作者: [Tianyi Zhang;Cuiyun Gao;Lei Ma;Michael R. Lyu;Miryung Kim]
通讯作者: Tianyi Zhang;Cuiyun Gao;Lei Ma;Michael R. Lyu;Miryung Kim
Software Engineering for Data Analytics
数据分析软件工程
DOI: 10.1109/ms.2020.2985775
发表时间: 2020
期刊: IEEE Software
影响因子: 3.3
作者: [Kim, Miryung]
通讯作者: Kim, Miryung
DOI: 10.48550/arxiv.2203.09615
发表时间: 2022-03
期刊:
影响因子: --
作者: [Chenxi Wang;Yifan Qiao;Haoran Ma;Shiafun Liu;Yiying Zhang;Wenguang Chen;R. Netravali;Miryung Kim;Guoqing Harry Xu]
通讯作者: Chenxi Wang;Yifan Qiao;Haoran Ma;Shiafun Liu;Yiying Zhang;Wenguang Chen;R. Netravali;Miryung Kim;Guoqing Harry Xu
Sibylvariant Transformations for Robust Text Classification
用于稳健文本分类的 Sibylvariant 变换
DOI: 10.18653/v1/2022.findings-acl.140
发表时间: 2022
期刊: Findings of the Association for Computational Linguistics: ACL 2022
影响因子: --
作者: [Harel-Canada, Fabrice, Gulzar, Muhammad Ali, Peng, Nanyun, Kim, Miryung]
通讯作者: Kim, Miryung
共 15 条
    Collaborative Research: SHF: Medium: Reinventing Fuzz Testing for Data and Compute Intensive Systems
    CHS: Medium: Collaborative Research: Code demography: Addressing information needs at scale for programming interface users and designers
    I-Corps: Interactive and Automated Debugging for Big Data Analytics
    SHF: Small: Analytical Support for Investigating Software Modifications in Collaborative Development Environment
    海外基金