Using Off-Line Features and Synthetic Data for On-Line Handwritten Math Symbol Recognition

Using Off-Line Features and Synthetic Data for On-Line Handwritten Math Symbol Recognition
复制标题

使用离线特征和合成数据进行在线手写数学符号识别

DOI:
10.1109/icfhr.2014.61
复制
发表时间:
2014
期刊:
2014 14th International Conference on Frontiers in Handwriting Recognition
影响因子:
--
通讯作者:
R. Zanibbi
R. Zanibbi
中科院分区:
--
文献类型:
--
作者:
Kenny Davila;S. Ludi;R. Zanibbi

文献摘要

被引文献

相似文献

提出了一种基于离线特征自适应和合成数据生成的手写数学符号在线识别方法。我们使用四种不同的分类方法来比较我们的方法的性能:AdaBoost。具有C4.5决策树、随机森林和具有线性和高斯核的支持向量机的M1。尽管计时信息可以从在线数据中提取,但我们的特征集基于形状描述,以更好地容忍绘制过程的变化。我们的主要数据集来自2012年和2013年的在线手写数学表达式识别竞赛(CROHME)。CROHME数据集中的类表示偏差通过使用弹性失真模型为未被表示的类生成样本来减轻。我们的结果表明,为未被充分代表的类生成合成数据可能会导致每类平均准确率的提高。我们还使用Math Brush数据集对我们的系统进行了测试,获得了89.87%的TOP-1准确率,这与最近在同一数据集上发布的其他方法的最佳结果相当。
We present an approach for on-line recognition of handwritten math symbols using adaptations of off-line features and synthetic data generation. We compare the performance of our approach using four different classification methods: AdaBoost. M1 with C4.5 decision trees, Random Forests and Support-Vector Machines with linear and Gaussian kernels. Despite the fact that timing information can be extracted from on-line data, our feature set is based on shape description for greater tolerance to variations of the drawing process. Our main datasets come from the Competition on Recognition of Online Handwritten Mathematical Expressions (CROHME) 2012 and 2013. Class representation bias in CROHME datasets is mitigated by generating samples for underrepresented classes using an elastic distortion model. Our results show that generation of synthetic data for underrepresented classes might lead to improvements of the average per-class accuracy. We also tested our system using the Math Brush dataset achieving a top-1 accuracy of 89.87% which is comparable with the best results of other recently published approaches on the same dataset.