Feedback is Needed for Retakes: An Explainable Poor Image Notification Framework for the Visually Impaired

Feedback is Needed for Retakes: An Explainable Poor Image Notification Framework for the Visually Impaired
复制标题

DOI:
10.1109/honet56683.2022.10019010
复制
发表时间:
2022-11
期刊:
2022 IEEE 19th International Conference on Smart Communities: Improving Quality of Life Using ICT, IoT and AI (HONET)
影响因子:
--
通讯作者:
Kazuya Ohata;Shunsuke Kitada;H. Iyatomi
Kazuya Ohata;Shunsuke Kitada;H. Iyatomi
中科院分区:
其他
文献类型:
--
作者:
Kazuya Ohata;Shunsuke Kitada;H. Iyatomi

文献摘要

相似文献

我们提出了一个简单而有效的图像字幕框架,可以确定图像的质量,并通知用户的原因,在图像中的任何缺陷。我们的框架首先确定图像的质量,然后只使用那些被确定为高质量的图像生成字幕。如果图像质量低,则缺陷特征通知用户重新拍摄,并且重复该循环,直到输入图像被认为具有高质量。作为框架的一个组成部分,我们训练和评估了一个低质量图像检测模型,该模型同时学习识别图像和单个缺陷的难度,我们证明了我们的建议可以用足够的分数解释缺陷的原因。我们还评估了一个数据集,其中低质量图像被我们的框架移除,并发现所有四个常见指标的值都有所改善(例如,BLEU-4、METEOR、ROUGE-L、CIDER),证实了通用图像字幕能力的改进。我们的框架将帮助视障人士,谁有困难判断图像质量。
We propose a simple yet effective image captioning framework that can determine the quality of an image and notify the user of the reasons for any flaws in the image. Our framework first determines the quality of images and then generates captions using only those images that are determined to be of high quality. The user is notified by the flaws feature to retake if image quality is low, and this cycle is repeated until the input image is deemed to be of high quality. As a component of the framework, we trained and evaluated a low-quality image detection model that simultaneously learns difficulty in recognizing images and individual flaws, and we demonstrated that our proposal can explain the reasons for flaws with a sufficient score. We also evaluated a dataset with low-quality images removed by our framework and found improved values for all four common metrics (e.g., BLEU-4, METEOR, ROUGE-L, CIDEr), confirming an improvement in general-purpose image captioning capability. Our framework would assist the visually impaired, who have difficulty judging image quality.