CAREER: Principled yet practical observability for a microservices-based cloud
CAREER: Principled yet practical observability for a microservices-based cloud
批准号:
2340128
负责人:
Raja Sambasivan
金额:
$60.95万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2024
资助国家:
美国
项目状态:
未结题
起止时间:
2024-07-01 至 2029-06-30
中文摘要
社会在日常生活的几乎每个方面都依赖于使用微服务架构构建的基于云的软件服务-例如,购物看电影工作 虽然微服务架构有很多优点,但它有一个关键的缺点。 观察用户如何请求(例如,购买一本书)是极具挑战性的,因为它们涉及许多更简单(微)服务之间的无数交互。 这种可观察性的缺乏使重要的管理任务复杂化,例如问题诊断和资源管理。 以前的研究已经证明了分布式跟踪的强大潜力--它捕获微服务如何交互以处理请求的图形--以提供微服务的可观察性。 但是,在现实世界中的结果令人失望。 这种潜在和现实之间的差距之所以出现,是因为研究工作假设了捕捉各种行为并且没有数据丢失的原则性跟踪图。 但是,在实践中,服务从来没有得到很好的检测,数据丢失是常见的。 该提案的总体目标是创建一个新的跟踪平台,该平台可以自动推断出使跟踪具有原则性所需的数据。这样做将实现分布式跟踪在微服务可观察性方面的巨大潜力,提高现有基于跟踪的管理工具的实用性,并启用变革性的新工具。这些成果将提高社会所依赖的软件服务的弹性和效率。该项目的见解将为大学课程中有关调试云系统的适龄课程材料和项目、高中研究项目和中学外展项目提供信息。该项目提出了一种新颖的跟踪平台,该平台使用两组原语自动丰富基于跨度的跟踪。 1)happens-before并发/等待原语和2)漏洞和漏洞覆盖原语。 前者允许识别请求的关键路径,支持松弛分析、有针对性的性能调试和精确的资源分配决策。 后者允许数据丢失的区域以跟踪沿着表示,并预测其中可能执行的工作。 由于调度决策可能会混淆因果结构,该项目将研究主动探测方法来梳理并发和等待关系。 为了解释不确定性,它将探索概率数据模型来表示原语。 该项目将通过修改现有的自动缩放解决方案和性能调试工具来使用它们来展示原语的价值。 它还将展示一种新的迹子图采样方法,使孔和孔覆盖原语成为可能。 该奖项反映了NSF的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Society relies on cloud-based software services built using the microservices architecture in almost every aspect of their everyday lives---e.g., to shop, watch movies, and work. Though the microservices architecture has many advantages, it has one critical drawback. Observing how user requests (e.g., to buy a book) are processed by services is extremely challenging because they involve myriad interactions among many simpler (micro)services. This lack of observability complicates important management tasks, such as problem diagnosis and resource management. Previous research has demonstrated the strong potential of distributed tracing---which captures graphs of how microservices interact to process requests---to provide microservice observability. But, results in real-world settings have been disappointing. This gap between potential and reality occurs because research efforts assume principled trace graphs that capture a variety of behaviors and have no data loss. But, in practice, services are never well-instrumented and data loss is common. The overarching goal of this proposal is to create a new tracing platform that automatically infers the data needed to make traces principled. Doing so will actualize distributed tracings' vast potential for microservice observability, improve the utility of existing tracing-based management tools, and enable transformative new tools. These outcomes will improve the resiliency and efficiency of the software services society depends on. Insights from this project will inform age-appropriate course material and projects in a college course on debugging cloud systems, a high-school research program, and a middle-school outreach program.This project proposes a novel tracing platform that automatically enriches span-based traces with two sets of primitives. 1) The happens-before concurrency/wait primitives and 2) the holes and holes covering primitives. The former allows requests' critical paths to be identified, enabling slack analyses, targeted performance debugging, and precise resource allocation decisions. The latter allows areas of data loss to be expressed in traces along with predictions of what work might execute in them. Since scheduling decisions may obfuscate causal structure, the project will investigate active probing methods to tease out concurrent and waiting relationships. To account for uncertainty, it will explore probabilistic data models to represent the primitives. The project will demonstrate the value of the primitives by modifying an existing auto-scaling solution and performance-debugging tool to use them. It will also demonstrate a new trace subgraph sampling approach made possible by the holes and holes covering primitive. The proposed platform, improved management tools, and inference methods will be publicly available to benefit the computer science and microservice observability communities.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CSR: Small: A Just-in-Time, Cross-Layer Instrumentation Framework for Diagnosing Performance Problems in Distributed Applications
-
批准号:2016178
-
项目类别:Standard Grant
-
资助金额:$39.79万
-
财政年份:2019
-
负责人:Raja Sambasivan
-
依托单位:
CSR: Small: A Just-in-Time, Cross-Layer Instrumentation Framework for Diagnosing Performance Problems in Distributed Applications
-
批准号:1815323
-
项目类别:Standard Grant
-
资助金额:$46.02万
-
财政年份:2018
-
负责人:Raja Sambasivan
-
依托单位:
海外基金