CODA: Toward Automatically Identifying and Scheduling Coflows in the Dark

CODA: Toward Automatically Identifying and Scheduling Coflows in the Dark
复制标题

DOI:
10.1145/2934872.2934880
复制
发表时间:
2016-08
期刊:
Proceedings of the 2016 ACM SIGCOMM Conference
影响因子:
--
通讯作者:
Hong Zhang;Li Chen;Bairen Yi;Kai Chen;Mosharaf Chowdhury;Yanhui Geng
Hong Zhang;Li Chen;Bairen Yi;Kai Chen;Mosharaf Chowdhury;Yanhui Geng
中科院分区:
其他
文献类型:
--
作者:
Hong Zhang;Li Chen;Bairen Yi;Kai Chen;Mosharaf Chowdhury;Yanhui Geng

文献摘要

被引文献

相似文献

最近的研究表明,使用coflows来利用应用程序级要求可以提高数据并行集群中的应用程序级通信性能。然而,现有的基于coflow的解决方案依赖于修改应用程序来提取coflow,使得它们不适用于许多实际场景。在本文中,我们提出了CODA,第一次尝试自动识别和调度coflows没有任何应用程序级的修改。我们采用增量聚类算法进行快速,应用程序透明的coflow识别和补充,提出了一个容错coflow调度,以减轻偶尔的识别错误。测试床实验和大规模模拟与生产工作负载表明,CODA可以识别coflow超过90%的准确性,其调度器是鲁棒的不准确性,使通信阶段完成2.4倍(5.1倍)平均快(95百分位数)相比,每流机制。总的来说,CODA的性能与需要修改应用程序的解决方案相当。
Leveraging application-level requirements using coflows has recently been shown to improve application-level communication performance in data-parallel clusters. However, existing coflow-based solutions rely on modifying applications to extract coflows, making them inapplicable to many practical scenarios. In this paper, we present CODA, a first attempt at automatically identifying and scheduling coflows without any application-level modifications. We employ an incremental clustering algorithm to perform fast, application-transparent coflow identification and complement it by proposing an error-tolerant coflow scheduler to mitigate occasional identification errors. Testbed experiments and large-scale simulations with production workloads show that CODA can identify coflows with over 90% accuracy, and its scheduler is robust to inaccuracies, enabling communication stages to complete 2.4x (5.1x) faster on average (95-th percentile) compared to per-flow mechanisms. Overall, CODA's performance is comparable to that of solutions requiring application modifications.