Bytewise Approximate Matching: The Good, The Bad, and The Unknown

Bytewise Approximate Matching: The Good, The Bad, and The Unknown
复制标题

字节近似匹配:好的、坏的和未知的

DOI:
10.15394/jdfsl.2016.1379
复制
发表时间:
2016
期刊:
J. Digit. Forensics Secur. Law
影响因子:
--
通讯作者:
I. Baggili
I. Baggili
中科院分区:
--
文献类型:
--
作者:
Vikram S. Harichandran;F. Breitinger;I. Baggili

文献摘要

被引文献

相似文献

散列函数在数字取证中被建立并且是众所周知的,其中它们通常用于证明完整性和文件标识(即,散列被查获设备上的所有文件,并将指纹与参考数据库进行比较)。然而,对于后一种操作,主动攻击者可以很容易地克服这种方法,因为传统的哈希被设计为对改变输入敏感;如果单个比特被翻转,输出将显著改变。因此,研究人员开发了近似匹配,这是一个相当新的,不太突出的领域,但被认为是传统哈希的更强大的对应物。自从近似匹配的概念提出以来,社区已经为这项技术构建了许多算法,扩展和其他应用程序,并且仍在研究新的概念以改善现状。在这篇调查文章中,我们从非技术角度对现有文献进行了高层次的回顾,并总结了近似匹配中的现有知识体系,特别关注字节算法。我们的贡献使研究人员和从业人员能够获得近似匹配的最新技术水平的概述,以便他们可以了解该领域的能力和挑战。简单地说,我们提出了术语,用例,分类,需求,测试方法,算法,应用程序,以及一系列主要和次要文献。
Hash functions are established and well-known in digital forensics, where they are commonly used for proving integrity and file identification (i.e., hash all files on a seized device and compare the fingerprints against a reference database). However, with respect to the latter operation, an active adversary can easily overcome this approach because traditional hashes are designed to be sensitive to altering an input; output will significantly change if a single bit is flipped. Therefore, researchers developed approximate matching, which is a rather new, less prominent area but was conceived as a more robust counterpart to traditional hashing. Since the conception of approximate matching, the community has constructed numerous algorithms, extensions, and additional applications for this technology, and are still working on novel concepts to improve the status quo. In this survey article, we conduct a high-level review of the existing literature from a non-technical perspective and summarize the existing body of knowledge in approximate matching, with special focus on bytewise algorithms. Our contribution allows researchers and practitioners to receive an overview of the state of the art of approximate matching so that they may understand the capabilities and challenges of the field. Simply, we present the terminology, use cases, classification, requirements, testing methods, algorithms, applications, and a list of primary and secondary literature.