SuSi: A Tool for the Fully Automated Classification and Categorization of Android Sources and Sinks

SuSi: A Tool for the Fully Automated Classification and Categorization of Android Sources and Sinks
复制标题

DOI:
--
复制
发表时间:
2013-05
期刊:
--
影响因子:
--
通讯作者:
Steven Arzt;Siegfried Rasthofer;E. Bodden
Steven Arzt;Siegfried Rasthofer;E. Bodden
中科院分区:
其他
文献类型:
--
作者:
Steven Arzt;Siegfried Rasthofer;E. Bodden

文献摘要

被引文献

相似文献

当今的智能手机用户面临安全困境:他们安装的许多应用程序都以隐私敏感的数据运行,尽管他们可能源于开发人员,他们的可信度很难判断。但是,这些工具的行为与他们配置的隐私政策一样好。敏感的数据以及可能泄漏到不信任的观察者的水槽我们表明,至少对于Android而言,API包括数百个来源和水槽。直接从我们的培训集中汇入Android源代码。水槽(例如,网络,文件等),平均精度和召回约89%。源和水槽的列表在很大程度上是不完整的,因此允许许多潜在的数据泄漏。
Today’s smartphone users face a security dilemma: many apps they install operate on privacy-sensitive data, although they might originate from developers whose trustworthiness is hard to judge. Researchers have proposed more and more sophisticated static and dynamic analysis tools as an aid to assess the behavior of such applications. Those tools, however, are only as good as the privacy policies they are configured with. Policies typically refer to a list of sources of sensitive data as well as sinks which might leak data to untrusted observers. Sources and sinks are a moving target: new versions of the mobile operating system regularly introduce new methods, and security tools need to be reconfigured to take them into account. In this work we show that, at least for the case of Android, the API comprises hundreds of sources and sinks. We propose SuSi, a novel and fully automated machine-learning approach for identifying sources and sinks directly from the Android source code. On our training set, SuSi achieves a recall and precision of more than 92%. To provide more fine-grained information, SuSi further categorizes the sources (e.g., unique identifier, location information, etc.) and sinks (e.g., network, file, etc.), with an average precision and recall of about 89%. We also show that many current program analysis tools can be circumvented because they use hand-picked lists of source and sinks which are largely incomplete, hence allowing many potential data leaks to go unnoticed.