BCFA: Bespoke Control Flow Analysis for CFA at Scale

BCFA: Bespoke Control Flow Analysis for CFA at Scale
复制标题

DOI:
10.1145/3377811.3380435
复制
发表时间:
2020-05
期刊:
2020 IEEE/ACM 42nd International Conference on Software Engineering (ICSE)
影响因子:
--
通讯作者:
Ramanathan Ramu;Ganesha Upadhyaya;H. Nguyen;Hridesh Rajan
Ramanathan Ramu;Ganesha Upadhyaya;H. Nguyen;Hridesh Rajan
中科院分区:
其他
文献类型:
--
作者:
Ramanathan Ramu;Ganesha Upadhyaya;H. Nguyen;Hridesh Rajan

文献摘要

被引文献

相似文献

许多数据驱动的软件工程任务,如发现编程模式、挖掘API规范等,在控制流图(cfg)上大规模地执行源代码分析。分析数百万个CFG可能是昂贵的,并且分析的性能在很大程度上取决于底层CFG遍历策略。最先进的分析框架使用固定的遍历策略。我们认为单一遍历策略不适合所有类型的分析和cfg,并提出了定制控制流分析(BCFA)。给定一个控制流分析(CFA)和大量的CFG, BCFA为每个CFG选择最有效的遍历策略。BCFA通过分析CFA代码提取CFA的一组属性,并将其与CFG的分支因子、循环度等属性相结合,选择最优的遍历策略。我们已经在Boa实现了BCFA,并使用一组代表性的静态分析来评估BCFA,这些分析主要涉及遍历cfg和两个包含28.7万个和1.62亿个cfg的大型数据集。结果表明,BCFA可以将大规模分析的速度提高1% ~ 28%。此外,BCFA的开销很低;小于0.2%,误判率低;小于0.01%。
Many data-driven software engineering tasks such as discovering programming patterns, mining API specifications, etc., perform source code analysis over control flow graphs (CFGs) at scale. Analyzing millions of CFGs can be expensive and performance of the analysis heavily depends on the underlying CFG traversal strategy. State-of-the-art analysis frameworks use a fixed traversal strategy. We argue that a single traversal strategy does not fit all kinds of analyses and CFGs and propose bespoke control flow analysis (BCFA). Given a control flow analysis (CFA) and a large number of CFGs, BCFA selects the most efficient traversal strategy for each CFG. BCFA extracts a set of properties of the CFA by analyzing the code of the CFA and combines it with properties of the CFG, such as branching factor and cyclicity, for selecting the optimal traversal strategy. We have implemented BCFA in Boa, and evaluated BCFA using a set of representative static analyses that mainly involve traversing CFGs and two large datasets containing 287 thousand and 162 million CFGs. Our results show that BCFA can speedup the large scale analyses by 1%-28%. Further, BCFA has low overheads; less than 0.2%, and low misprediction rate; less than 0.01%.