Deterministic Routing between Layout Abstractions for Multi-Scale Classification of Visually Rich Documents

Deterministic Routing between Layout Abstractions for Multi-Scale Classification of Visually Rich Documents
复制标题

DOI:
10.24963/ijcai.2019/466
复制
发表时间:
2019-08
期刊:
--
影响因子:
--
通讯作者:
Ritesh Sarkhel;Arnab Nandi
Ritesh Sarkhel;Arnab Nandi
中科院分区:
其他
文献类型:
--
作者:
Ritesh Sarkhel;Arnab Nandi

文献摘要

相似文献

对视觉丰富的异类文档进行分类是一项具有挑战性的任务。如果允许的最大推理周转时间受到阈值的限制,则此任务的难度会更大。推理成本的增加,与分类能力的有限收益相比,使得当前的多尺度方法在这种情况下不可行。这项工作有两个主要贡献。首先,我们提出了一种空间金字塔模型,通过利用文档布局的内在层次性,从视觉丰富的文档中提取具有高度区分性的多尺度特征描述符。其次,利用空间金字塔模型,提出了一种加速端到端推理的确定性路由方案。我们开发了一个沿深度方向可分离的多列卷积网络来实现我们的方法。我们在四个公开可用的视觉丰富文档的基准数据集上对所提出的方法进行了评估。结果表明,与现有的分类方法相比,本文提出的方法在分类准确率和总的推理率方面都表现出了较好的性能。
Classifying heterogeneous visually rich documents is a challenging task. Difficulty of this task increases even more if the maximum allowed inference turnaround time is constrained by a threshold. The increased overhead in inference cost, compared to the limited gain in classification capabilities make current multi-scale approaches infeasible in such scenarios. There are two major contributions of this work. First, we propose a spatial pyramid model to extract highly discriminative multi-scale feature descriptors from a visually rich document by leveraging the inherent hierarchy of its layout. Second, we propose a deterministic routing scheme for accelerating end-to-end inference by utilizing the spatial pyramid model. A depth-wise separable multi-column convolutional network is developed to enable our method. We evaluated the proposed approach on four publicly available, benchmark datasets of visually rich documents. Results suggest that our proposed approach demonstrates robust performance compared to the state-of-the-art methods in both classification accuracy and total inference turnaround.