FilteredWeb: A framework for the automated search-based discovery of blocked URLs

FilteredWeb: A framework for the automated search-based discovery of blocked URLs
复制标题

FilteredWeb:基于搜索自动发现被阻止 URL 的框架

DOI:
--
复制
发表时间:
2017
期刊:
Traffic Monitoring and Analysis
影响因子:
--
通讯作者:
Joss Wright
Joss Wright
中科院分区:
--
文献类型:
--
作者:
Alexander Darer;Oliver Farnan;Joss Wright

文献摘要

被引文献

相似文献

已经提出了各种方法来创建和维护可能被过滤的URL的列表,以允许测量世界各地正在进行的互联网审查。虽然测试已知资源的过滤证据可能相对简单,但如果给定适当的Vantage,发现先前未知的过滤Web资源仍然是一个开放的挑战。我们提出了一个新的框架,自动化的过程中发现过滤资源,通过使用自适应查询知名的搜索引擎。我们的系统应用信息检索算法来隔离已知的过滤网页中的特征语言模式,这些被用作网络搜索查询的基础。这些搜索的结果URL将被检查是否存在过滤证据,新发现的被阻止资源将被反馈到系统中,以检测进一步过滤的内容。我们实施这个框架,适用于中国作为一个案例研究,表明该方法在检测大量以前未知的过滤网页,使正在进行的检测互联网过滤,因为它的发展作出了重大贡献是显而易见的有效。在部署时,截至2017年2月,该系统在中国发现了1355个中毒域名-比当时最广泛使用的已发布过滤器列表多30倍。其中,759个不在Alexa Top 1000域名列表中,这表明该框架能够找到更模糊的过滤内容。此外,我们对过滤URL的初步分析,以及用于发现它们的搜索词,进一步深入了解了目前在中国被屏蔽的内容的性质。
Various methods have been proposed for creating and maintaining lists of potentially filtered URLs to allow for measurement of ongoing internet censorship around the world. Whilst testing a known resource for evidence of filtering can be relatively simple, given appropriate vantage points, discovering previously unknown filtered web resources remains an open challenge. We present a novel framework for automating the process of discovering filtered resources through the use of adaptive queries to well-known search engines. Our system applies information retrieval algorithms to isolate characteristic linguistic patterns in known filtered web pages; these are used as the basis for web search queries. The resulting URLs of these searches are checked for evidence of filtering, and newly discovered blocked resources will be fed back into the system to detect further filtered content. Our implementation of this framework, applied to China as a case study, shows the approach is demonstrably effective at detecting significant numbers of previously unknown filtered web pages, making a significant contribution to the ongoing detection of internet filtering as it develops. When deployed, this system was used to discover 1355 poisoned domains within China as of Feb 2017 — 30 times more than in the most widely-used published filter list of the time. Of these, 759 are outside of the Alexa Top 1000 domains list, demonstrating the capability of this framework to find more obscure filtered content. Further, our initial analysis of filtered URLs, and the search terms that were used to discover them, gives further insight into the nature of the content currently being blocked in China.