Ballot Tabulation Using Deep Learning

Ballot Tabulation Using Deep Learning
复制标题

DOI:
10.1109/iri58017.2023.00026
复制
发表时间:
2023-08
期刊:
2023 IEEE 24th International Conference on Information Reuse and Integration for Data Science (IRI)
影响因子:
--
通讯作者:
Fei Zhao;Chengcui Zhang;Nitesh Saxena;D. Wallach;AKM SHAHARIAR AZAD RABBY
Fei Zhao;Chengcui Zhang;Nitesh Saxena;D. Wallach;AKM SHAHARIAR AZAD RABBY
中科院分区:
其他
文献类型:
--
作者:
Fei Zhao;Chengcui Zhang;Nitesh Saxena;D. Wallach;AKM SHAHARIAR AZAD RABBY

文献摘要

相似文献

目前部署的扫描和处理手写选票的选举系统不够复杂,无法处理未充分填写的标记(例如,部分填充),不正确的标记(例如,使用复选标记或十字而不是填充气泡),或气泡外部的标记,而不是设置阈值来检测气泡内部的像素是否足够暗和密集以被计数为投票。目前沿着这条线的工作仍然在很大程度上受到其自动化程度的限制,需要大量的人力进行注释和裁定。在这项研究中,我们提出了一种高度自动化的基于深度学习(DL)标记分割模型的选票制表助手,能够准确识别合法的选票标记。为了比较的目的,一个高度定制的传统的计算机视觉(T-CV)标记分割为基础的方法也已经开发出与DL为基础的制表,包括详细的讨论。我们在两个真实的选举数据集上进行的实验在选票列表上达到了99.984%的最高准确率。为了进一步增强我们的DL模型检测训练数据集中代表性不足的标记的能力,例如,不充分或不正确的填充标记,我们提出了一个连体网络架构,使我们的DL模型,利用一个手标记的选票图像和相应的空白模板图像之间的对比功能,以检测标记。不需要额外的数据收集,通过结合这种新颖的网络架构,我们的基于DL模型的制表方法不仅实现了更高的准确性得分,而且大大降低了整体假阴性率。
Currently deployed election systems that scan and process hand-marked ballots are not sophisticated enough to handle marks insufficiently filled in (e.g., partially filled-in), improper marks (e.g., using check marks or crosses instead of filling in bubbles), or marks outside of bubbles, other than setting a threshold to detect whether the pixels inside bubbles are dark and dense enough to be counted as a vote. The current works along this line are still largely limited by their degree of automation and require substantial manpower for annotation and adjudication. In this study, we propose a highly automated deep learning (DL) mark segmentation model-based ballot tabulation assistant able to accurately identify legitimate ballot marks. For comparison purposes, a highly customized traditional computer vision (T-CV) mark segmentation-based method has also been developed to compare with the DL-based tabulator, with a detailed discussion included. Our experiments conducted on two real election datasets achieved the highest accuracy of 99.984% on ballot tabulation. In order to further enhance our DL model’s capability of detecting the marks that are underrepresented in training datasets, e.g., insufficiently or improperly filled marks, we propose a Siamese network architecture that enables our DL model to exploit the contrasting features between a hand-marked ballot image and its corresponding blank template image to detect marks. Without the need for extra data collection, by incorporating this novel network architecture, our DL model-based tabulation method not only achieved a higher accuracy score but also substantially reduced the overall false negative rate.