A generative vision model that trains with high data efficiency and breaks text-based CAPTCHAs

A generative vision model that trains with high data efficiency and breaks text-based CAPTCHAs
复制标题

DOI:
10.1126/science.aag2612
复制
发表时间:
2017-12-08
期刊:
影响因子:
56.9
通讯作者:
Phoenix, D. Scott
Phoenix, D. Scott
中科院分区:
综合性期刊1区
文献类型:
--
作者:
George, Dileep;Lehrach, Wolfgang;Phoenix, D. Scott

文献摘要

被引文献

相似文献

从几个例子中学习并推广到明显不同的情况是人类视觉智能的能力,尚未被领先的机器学习模型所匹配。通过从系统神经科学中汲取灵感,我们引入了一个概率生成模型,其中基于消息传递的推理以统一的方式处理识别,分割和推理。该模型具有出色的泛化和遮挡推理能力,在具有挑战性的场景文本识别基准测试中优于深度神经网络,同时数据效率提高300倍。此外,该模型从根本上打破了现代基于文本的CAPTCHA(完全自动化的公共图灵测试,以区分计算机和人类)的防御,通过生成分割字符,而无需CAPTCHA特定的验证。我们的模型强调了数据效率和组合性等方面,这些方面在迈向通用人工智能的道路上可能很重要。
Learning from a few examples and generalizing to markedly different situations are capabilities of human visual intelligence that are yet to be matched by leading machine learning models. By drawing inspiration from systems neuroscience, we introduce a probabilistic generative model for vision in which message-passing-based inference handles recognition, segmentation, and reasoning in a unified way. The model demonstrates excellent generalization and occlusion-reasoning capabilities and outperforms deep neural networks on a challenging scene text recognition benchmark while being 300-fold more data efficient. In addition, the model fundamentally breaks the defense of modern text-based CAPTCHAs (Completely Automated Public Turing test to tell Computers and Humans Apart) by generatively segmenting characters without CAPTCHA-specific heuristics. Our model emphasizes aspects such as data efficiency and compositionality that may be important in the path toward general artificial intelligence.