Fingerprinting Keywords in Search Queries over Tor

Fingerprinting Keywords in Search Queries over Tor
复制标题

对 Tor 搜索查询中的关键字进行指纹识别

DOI:
--
复制
发表时间:
2017
影响因子:
--
通讯作者:
Nicholas Hopper
Nicholas Hopper
中科院分区:
--
文献类型:
--
作者:
Se Eun Oh;Shuai Li;Nicholas Hopper

文献摘要

被引文献

相似文献

摘要搜索引擎查询包含大量关于用户的隐私和潜在危害信息。防止搜索引擎识别查询源以及互联网服务提供商(ISP)识别查询内容的一种技术是在诸如ToR的匿名网络上查询搜索引擎。在本文中,我们研究了网站指纹识别可以扩展到对Web应用程序的单个查询或关键字进行指纹识别的程度,我们称之为关键字指纹识别(KF)。我们表明,通过使用新的特定于任务的特征集的两阶段方法来增强流量分析,被动的网络对手在许多情况下可以击败使用Tor来保护搜索引擎查询的使用。我们探索了三个流行的搜索引擎,Google,Bing和DuckDuckGo,以及几种机器学习技术和各种实验场景。我们的实验结果表明,KF能够识别包含300个目标关键词中的一个的Google查询,召回率为80%,准确率为91%,而识别300个搜索关键词中的特定监控关键词的准确率为48%。我们还进一步调查了影响关键字指纹的因素,以了解搜索引擎和用户可能如何防止KF。
Abstract Search engine queries contain a great deal of private and potentially compromising information about users. One technique to prevent search engines from identifying the source of a query, and Internet service providers (ISPs) from identifying the contents of queries is to query the search engine over an anonymous network such as Tor. In this paper, we study the extent to which Website Fingerprinting can be extended to fingerprint individual queries or keywords to web applications, a task we call Keyword Fingerprinting (KF). We show that by augmenting traffic analysis using a two-stage approach with new task-specific feature sets, a passive network adversary can in many cases defeat the use of Tor to protect search engine queries. We explore three popular search engines, Google, Bing, and Duckduckgo, and several machine learning techniques with various experimental scenarios. Our experimental results show that KF can identify Google queries containing one of 300 targeted keywords with recall of 80% and precision of 91%, while identifying the specific monitored keyword among 300 search keywords with accuracy 48%. We also further investigate the factors that contribute to keyword fingerprintability to understand how search engines and users might protect against KF.