Robust Performance Metrics for Authentication Systems

Robust Performance Metrics for Authentication Systems
复制标题

DOI:
10.14722/ndss.2019.23351
复制
发表时间:
2019
期刊:
Proceedings 2019 Network and Distributed System Security Symposium
影响因子:
--
通讯作者:
Shridatt Sugrim;Can Liu;Meghan McLean;J. Lindqvist
Shridatt Sugrim;Can Liu;Meghan McLean;J. Lindqvist
中科院分区:
其他
文献类型:
--
作者:
Shridatt Sugrim;Can Liu;Meghan McLean;J. Lindqvist

文献摘要

相似文献

研究已经产生了多种使用机器学习的身份验证系统。然而,没有一致的方法来报告绩效指标,并且报告的指标不充分。在这项工作中,我们证明了用于报告绩效的几个常见指标,例如最大准确度 (ACC)、等错误率 (EER) 和 ROC 曲线下面积 (AUROC),本质上是有缺陷的。这些通用指标隐藏了系统在实施时必须做出的固有权衡的细节。我们的研究结果表明,当前的指标无法深入了解系统性能在其设计的理想条件之外如何降低。我们认为,必须提供充分的绩效报告才能进行有意义的评估,而当前常用的方法在这方面失败了。我们提出了非标准化分数频率计数 (FCS),以证明导致这些失败的数学基础,并展示如何避免它们。 FCS 可用于增强性能报告,以便以可视化方式进行系统间比较。当与接收者操作特征曲线 (ROC) 一起报告时,这两个指标为当前报告指标的局限性提供了解决方案。最后,我们展示如何使用 FCS 和 ROC 指标来评估和比较不同的身份验证系统。
Research has produced many types of authentication systems that use machine learning. However, there is no consistent approach for reporting performance metrics and the reported metrics are inadequate. In this work, we show that several of the common metrics used for reporting performance, such as maximum accuracy (ACC), equal error rate (EER) and area under the ROC curve (AUROC), are inherently flawed. These common metrics hide the details of the inherent tradeoffs a system must make when implemented. Our findings show that current metrics give no insight into how system performance degrades outside the ideal conditions in which they were designed. We argue that adequate performance reporting must be provided to enable meaningful evaluation and that current, commonly used approaches fail in this regard. We present the unnormalized frequency count of scores (FCS) to demonstrate the mathematical underpinnings that lead to these failures and show how they can be avoided. The FCS can be used to augment the performance reporting to enable comparison across systems in a visual way. When reported with the Receiver Operating Characteristics curve (ROC), these two metrics provide a solution to the limitations of currently reported metrics. Finally, we show how to use the FCS and ROC metrics to evaluate and compare different authentication systems.