The impact of tool configuration spaces on the evaluation of configurable taint analysis for Android

The impact of tool configuration spaces on the evaluation of configurable taint analysis for Android
复制标题

DOI:
10.1145/3460319.3464823
复制
发表时间:
2021-07
期刊:
Proceedings of the 30th ACM SIGSOFT International Symposium on Software Testing and Analysis
影响因子:
--
通讯作者:
Austin Mordahl;Shiyi Wei
Austin Mordahl;Shiyi Wei
中科院分区:
其他
文献类型:
--
作者:
Austin Mordahl;Shiyi Wei

文献摘要

被引文献

相似文献

最受欢迎的静态污染分析工具允许用户通过配置选项更改基础分析算法。但是,较大的配置空间使开发人员和用户都难以了解这些工具的全部功能,并且迄今为止的研究仅集中在各个配置上。在这项工作中,我们介绍了第一个评估Android污染分析工具中配置的研究,重点介绍了FlowDroid和DroidSafe的两个最受欢迎的工具。首先,我们执行手动代码调查,以更好地了解两个工具中如何实现配置。我们根据精确和健全的部分订单对配置选项设置的预期效果进行形式化,我们用来系统地测试配置空间。其次,我们创建了一个新的数据集,其中包括18个开源现实世界应用程序中的756个手动分类流量,并在此数据集和微基准测试上进行大规模实验。我们观察到,配置在两个工具的性能,精度和健全性上都取决于重大的权衡。迄今为止的研究将得出有关工具功能的不同结论,即他们考虑配置或使用现实世界数据集。此外,我们通过统计分析来研究各个选项,并为用户提出可行的建议,使用户将工具调整为自己目的。最后,我们使用部分订单来测试工具配置空间并检测21个实例,其中选项以意外和不正确的方式行事,证明了对配置空间进行严格测试的需求。
The most popular static taint analysis tools for Android allow users to change the underlying analysis algorithms through configuration options. However, the large configuration spaces make it difficult for developers and users alike to understand the full capabilities of these tools, and studies to-date have only focused on individual configurations. In this work, we present the first study that evaluates the configurations in Android taint analysis tools, focusing on the two most popular tools, FlowDroid and DroidSafe. First, we perform a manual code investigation to better understand how configurations are implemented in both tools. We formalize the expected effects of configuration option settings in terms of precision and soundness partial orders which we use to systematically test the configuration space. Second, we create a new dataset of 756 manually classified flows across 18 open-source real-world apps and conduct large-scale experiments on this dataset and micro-benchmarks. We observe that configurations make significant tradeoffs on the performance, precision, and soundness of both tools. The studies to-date would reach different conclusions on the tools' capabilities were they to consider configurations or use real-world datasets. In addition, we study the individual options through a statistical analysis and make actionable recommendations for users to tune the tools to their own ends. Finally, we use the partial orders to test the tool configuration spaces and detect 21 instances where options behaved in unexpected and incorrect ways, demonstrating the need for rigorous testing of configuration spaces.