Limitless HTTP in an HTTPS World: Inferring the Semantics of the HTTPS Protocol without Decryption

Limitless HTTP in an HTTPS World: Inferring the Semantics of the HTTPS Protocol without Decryption
复制标题

DOI:
10.1145/3292006.3300025
复制
发表时间:
2018-05
期刊:
Proceedings of the Ninth ACM Conference on Data and Application Security and Privacy
影响因子:
--
通讯作者:
Blake Anderson;A. Chi;Scott Dunlop;D. McGrew
Blake Anderson;A. Chi;Scott Dunlop;D. McGrew
中科院分区:
其他
文献类型:
--
作者:
Blake Anderson;A. Chi;Scott Dunlop;D. McGrew

文献摘要

被引文献

相似文献

我们提出了一种新的分析技术,用于从对HTTPS的被动观察中推断HTTP语义,该技术可以推断重要字段的值,包括状态代码、内容类型和服务器,以及几个额外的HTTP报头字段的存在或不存在,例如Cookie和Referer。我们的目标是提高对HTTPS保密限制的理解,并探索流量分析的良性应用,在某些场景中可以取代HTTPS拦截和静态私钥。我们发现,我们的技术提高了恶意软件检测的效率,但它们并不能实现针对Tor的更强大的网站指纹攻击。我们的一组更广泛的结果引发了人们对TLS相对于用户隐私期望的保密目标的担忧,这证明了未来的研究。我们将我们的方法应用于从实验室环境中自动运行的Firefox 58.0、Chrome 63.0和Tor Browser 7.0.11以及从恶意软件沙箱中运行的应用程序收集的数据的HTTP/1.1和HTTP/2的语义。我们通过从执行后的RAM中提取解密所需的密钥材料,从恶意软件沙箱中获取各种应用程序的基本事实明文。我们开发了一种迭代方法来同时解决多类(字段值)和二进制(字段存在)分类问题,并且我们证明了我们的推理算法对于大多数被检查的Http字段获得了大于0.900的未加权$F1$分数。
We present new analytic techniques for inferring HTTP semantics from passive observations of HTTPS that can infer the value of important fields including the status-code, Content-Type, and Server, and the presence or absence of several additional HTTP header fields, e.g., Cookie and Referer. Our goals are to improve the understanding of the confidentiality limitations of HTTPS, and to explore benign uses of traffic analysis that could replace HTTPS interception and static private keys in some scenarios. We found that our techniques increase the efficacy of malware detection, but they do not enable more powerful website fingerprinting attacks against Tor. Our broader set of results raises concerns about the confidentiality goals of TLS relative to a user's expectation of privacy, warranting future research. We apply our methods to the semantics of both HTTP/1.1 and HTTP/2 on data collected from automated runs of Firefox 58.0, Chrome 63.0, and Tor Browser 7.0.11 in a lab setting, and from applications running in a malware sandbox. We obtain ground truth plaintext for a diverse set of applications from the malware sandbox by extracting the key material needed for decryption from RAM post-execution. We developed an iterative approach to simultaneously solve several multi-class (field values) and binary (field presence) classification problems, and we show that our inference algorithm achieves an unweighted $F_1$ score greater than 0.900 for most HTTP fields examined.