课题基金 / 基金详情

Detecting Training Abuses in Neural Nets

Detecting Training Abuses in Neural Nets
检测神经网络中的训练滥用
批准号:
2301656
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2019
资助国家:
英国
项目状态:
已结题
起止时间:
2019 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
许多军事系统执行分类任务。例如,可能需要一个系统来区分盟军坦克和敌人坦克(这是机器学习中的一个经典问题)。现代机器学习方法正被用于军事领域内和更广泛的分类问题。神经网络是一种技术,粗略地说,它做出决策的方式类似于人脑的工作方式,它发挥着特别突出的作用。大多数工作都是在假设一切都是良性的基础上进行的。但想象一下,如果敌人想要让你的分类器工作得很好,除非你面临非常具体的分类任务。例如,一辆外观特殊的敌人坦克可以被归类为盟军坦克,后果非常严重。敌人能策划这样的行为吗?在某些情况下,是的!这取决于如何以及由谁构建分类器系统。这类系统的建造往往以某种方式外包,例如因为采购人缺乏设计有效系统的计算能力,或通过使用公共生成的组件。我们经常将隐藏的恶意功能称为“陷阱门”,这些功能可以在方便的时候调用。暗门通常很难被察觉。想象一下这样一个系统,它对你提供的数千个坦克实例进行了完美的分类。看起来这是一个非常好的系统。但该系统可能经过了训练,以至于一辆侧面涂有“666”的敌方坦克被错误归类。如果您不知道这个特定的条件,您就没有什么理由生成测试用例来发现它。众所周知,神经网络在呈现它们如何做出决策方面也是出了名的不透明,这使得这种暗门检测特别困难。我们可能会合理地问,我们能否或在多大程度上能够很好地探测到这样的暗门。可以在不同的层面上寻求理解。因此,确定系统中是否有活门(是/否)比寻找特定的活门条件(上面所示的“666”)更简单,也不那么雄心勃勃。虽然在文献中有相当多的关于活板门的文献,通常是关于种植或检测活板门的问题,但似乎很少关注如何描述它们。显然,任何检测技术在某些活门上都可能比在其他活门上更成功。然而,这就提出了一个问题,即如何描述这项技术在哪些方面效果很好,哪些地方效果不佳。本项目的主要目标是采用严格的检测方法,这需要对活板门有细微差别的理解。特别是,活板门的特征及其性质的测量,例如活板门实例与正常实例的偏差程度,是必不可少的。如果现在考虑活板门的生成,则活板门的特征允许对我们希望插入的活板门具有的属性进行更精细的规范。这有两个目的:第一,它促进了一种更微妙的世代能力,用于实际操作目的,即对于希望从现实世界中安装活动门中受益的人;第二,它允许研究人员(最初是我们自己!)以生成多组活板门,以严格评估检测技术。我们可以通过某种方式定义“覆盖”活动门空间意味着什么,就像我们在一般测试中覆盖输入或其他空间一样。由于没有现存的可操作的活板门特征,显然也没有现存的世代能力。
英文摘要
Many military systems carry out classification tasks. For example, a system might be required to distinguish between an allied tank and an enemy tank (a classic problem in machine learning). Modern machine learning approaches are being brought to bear on classification problems within the military domain and more widely. Neural networks, a technology that, loosely speaking, makes decisions in a manner analogous to the way the human brain works, play a particularly prominent role. Most work proceeds on the assumption that all is benign. But imagine if an enemy wanted to cause your classifier to work well, except when presented with a very specific classification task. For example, an enemy tank with a particular appearance could be classified as an allied one, with very significant consequences. Can an enemy engineer such behaviour? In certain circumstances, yes! It depends on how and by whom the classifier system was built. The building of such systems is often outsourced in some way, e.g. because the procurer lacks the computational capability to craft an effective system or by the use of publicly generated components. We often refer to hidden malicious functionality that can be invoked when convenient as a 'trapdoor'. Trapdoors are often very difficult to detect. Imagine a system that classified perfectly thousands of tank examples provided by you. It seems like this is a very good system. But the system may have been trained so that an enemy tank with "666" painted on its side is misclassified. If you don't know this specific condition you would have little reason to generate a test example to discover it. Neural networks are also notoriously opaque in rendering apparent how they make decisions and this makes this sort of trapdoor detection particularly hard. We might reasonably ask whether or how well we can detect such trapdoors. There are various levels at which understanding may be sought. Thus, determining whether a system has a trapdoor in it (yes/no) is a simpler and less ambitious task than seeking the specific trapdoor condition (the "666' indicated above). Though there is a fair amount on trapdoors in the literature, typically addressing issues of planting or detecting trapdoors, there appears to be little concerned with characterising them. It would seem clear that any detection technique is likely to be more successful on some trapdoors than on others. This raises the question, however, as to how to describe those where the technique works well and those where it performs less well. A rigorous approach to detection, the primary goal of this project, requires a nuanced understanding of trapdoors. In particular, a characterisation of trapdoors together with measurements of their properties, e.g. how much a trapdoor example deviates from a normal example, is essential. If trapdoor generation is now considered, the characterisation of trapdoors allows more refined specification of properties we would like an inserted trapdoor to have. This serves two purposes: firstly, it facilitates a more nuanced generational capability for practical operational purposes, i.e. for someone who wants to benefit from planting a trapdoor in the real world; and secondly, it allows researchers (initially ourselves!) to generate sets of trapdoors for rigorous evaluation of detection techniques. We can define what it means to 'cover' the trapdoor space in some way, much as we cover input or other space in general testing. Since there is no extant workable characterisation of trapdoors there is also clearly no extant generational capability.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金