Classification of Blind Users’ Image Exploratory Behaviors Using Spiking Neural Networks

Classification of Blind Users’ Image Exploratory Behaviors Using Spiking Neural Networks
复制标题

DOI:
10.1109/tnsre.2019.2959555
复制
发表时间:
2019-12
影响因子:
4.9
通讯作者:
Ting Zhang;B. Duerstock;J. Wachs
Ting Zhang;B. Duerstock;J. Wachs
中科院分区:
工程技术2区
文献类型:
--
作者:
Ting Zhang;B. Duerstock;J. Wachs

文献摘要

相似文献

盲人采用多种程序来触觉探索图像。自动识别和分类用户的探索行为是开发智能系统的第一步,可以帮助用户更有效地探索图像。在本文中,一个计算框架的开发,分类不同的程序使用的盲人用户在图像探索。从用户的运动轨迹中提取平移、旋转和尺度不变的特征。这些特征被分为数值和逻辑特征,并被送入神经网络。更具体地说,我们训练了尖峰神经网络(SNN),以进一步将数字特征编码为模型字符串。拟议的框架采用了基于距离的分类方案,以确定探索性程序的最终类别/标签。应用Dempster-Shafter理论(DST)对从所有特征获得的距离进行整合。通过对不同脉冲神经元动力学的实验,该框架取得了良好的性能,分类准确率为95.89%。它是非常有效的编码和分类时空数据,相比动态时间弯曲和隐马尔可夫模型的准确率为61.30%和28.70%。建议的框架作为智能接口的发展,提高盲人的图像探索经验的基本块。
Individuals who are blind adopt multiple procedures to tactually explore images. Automatically recognizing and classifying users’ exploration behaviors is the first step towards the development of an intelligent system that could assist users to explore images more efficiently. In this paper, a computational framework was developed to classify different procedures used by blind users during image exploration. Translation-, rotation- and scale-invariant features were extracted from the trajectories of users’ movements. These features were divided as numerical and logical features and were fed into neural networks. More specifically, we trained spiking neural networks (SNNs) to further encode the numerical features as model strings. The proposed framework employed a distance-based classification scheme to determine the final class/label of the exploratory procedures. Dempster-Shafter Theory (DST) was applied to integrate the distances obtained from all the features. Through the experiments of different dynamics of spiking neurons, the proposed framework achieved a good performance with 95.89% classification accuracy. It is extremely effective in encoding and classifying spatio-temporal data, as compared to Dynamic Time Warping and Hidden Markov Model with 61.30% and 28.70% accuracy. The proposed framework serves as the fundamental block for the development of intelligent interfaces, enhancing the image exploration experience for the blind.