Performance Evaluation of Apache Kafka – A Modern Platform for Real Time Data Streaming

Performance Evaluation of Apache Kafka – A Modern Platform for Real Time Data Streaming
复制标题

Apache Kafka 的性能评估——现代实时数据流平台

DOI:
10.1109/iciptm54933.2022.9754154
复制
发表时间:
2022
期刊:
2022 2nd International Conference on Innovative Practices in Technology and Management (ICIPTM)
影响因子:
--
通讯作者:
Shashank Sahu
Shashank Sahu
中科院分区:
--
文献类型:
--
作者:
Shubham Vyas;R. Tyagi;Charu Jain;Shashank Sahu

文献摘要

被引文献

相似文献

当代企业对数据的及时可用性要求越来越高。许多实时数据流工具和技术能够满足业务期望。ApacheKafka是一种强大的开源分布式可扩展技术,能够以良好的吞吐量和延迟实现实时数据流。在传统的批处理中,数据是按组或批处理的,但在流服务中,数据记录是单独处理的,并且存在连续和实时的数据处理流。一旦数据在来源处可用,Kafka就可以检测并实时将其传输到目标应用程序。在做了文献调查之后,观察到到目前为止,对于不同的卷以及不同的分区数量和轮询间隔的值所做的实验是不够的。本研究的目的是详细阐述ApacheKafka的实现并对其性能进行评估。这项研究将分析流媒体平台的关键绩效指标,并将提供有益的见解。这些见解将有助于在ApacheKafka中设计优化的应用程序。基于文献调查后发现的差距,已经针对生产者和消费者API(应用程序编程接口)进行了多次实验。使用ApacheZooKeeper配置Kafka有助于推动结果,这些结果以表格形式为不同的轮询间隔、卷和分区值捕获。对所有试运行的数据进行了进一步分析,以得出结果部分所述的结论。此研究提供了有关在更改卷时对ApacheKafka流的CPU(中央处理器)和内存利用率的有价值的见解,并阐述了关键配置更改时对流性能的影响。
Current generation businesses become more demanding on timely availability of data. Many real-time data streaming tools and technologies are capable to meet business expectations. Apache Kafka is one of the capable open-source distributed scalable technology that enables real-time data streaming with good throughput and latency. In traditional batch processing, data is getting processed in groups or batches but in streaming services, data records are handled separately and there is a flow of data processing that is continuous and real-time. Once Data is available at the source, Kafka can detect and stream it in real-time to the target application. After doing the literature survey it was observed that there are insufficient experiments have been done till now with a variety of volumes and with different values of the number of partitions and polling intervals. The purpose of this study is to elaborate on Apache Kafka implementation and evaluate its performance. This study will analyse key performance indicators for the streaming platform and will provide useful insights from it. These insights will help to design optimized applications in Apache Kafka. Based on gaps identified after the literature survey, multiple experiments have been conducted for the producer and consumer API (Application Programming interface). Configuration of Kafka with Apache Zookeeper helped to drive the results which are captured in tabular form for different values of polling intervals, volumes, and partitions. Data for all test runs have been analysed further to drive the conclusions as mentioned in the results section. This study provides valuable insights about the utilization of CPU (Central Processing Unit) and memory for Apache Kafka streaming on changing volumes, also elaborates the impacts on streaming performance when key configurations are getting changed.