Teaching Machines to Write Like Humans Using L-attributed Grammar

Teaching Machines to Write Like Humans Using L-attributed Grammar
复制标题

使用 L 属性语法教机器像人类一样写作

DOI:
10.1016/j.engappai.2020.103489
复制
发表时间:
--
影响因子:
8
通讯作者:
Cheng-Lin Liu
Cheng-Lin Liu
中科院分区:
计算机科学2区
文献类型:
--
作者:
Yunxue Shao;Cheng-Lin Liu

文献摘要

相似文献

阅读和写作对人类来说很容易。手写体的自动阅读已经研究了几十年。用于阅读任务的机器学习算法通常需要大量数据才能达到与人类相似的精度,但也很难获得足够有意义的数据。自动书写任务还没有得到广泛的研究。在本文中,我们通过告诉机器使用l属性语法书写每个字符的方法来教机器像教孩子一样写作。在提出的TMTW (Teaching Machines To Write)交互系统的帮助下,作为教师的人只需要提供零件和控制线的书写顺序。该系统能够自动感知控制线和部件之间的关系,并构建相应的语法。采用自顶向下派生和笔画生成方法,根据学习到的语法生成不同的字符。只要机器会写,就可以应用于机器人控制或自动读取任务的训练样本生成。使用MNIST和CASIA数据集验证了该系统在不同语言上的有效性。机器编写的样本用于训练网络,并在MNIST测试集上对其进行评估。每个数字平均只使用大约20个语法,测试错误率为1.23%。将生成的样本和手写样本一起作为训练集,可以将测试错误率降低到0.61%。在CASIA数据集上进行了类似的实验,结果表明该方法可以有效地生成具有复杂结构的字符。本文中使用的源代码和语法已在https://github.com/step123456789/TMTW上公开提供。
Reading and writing are easy for humans. The automatic reading of handwritten characters has been studied for several decades. Machine learning algorithms for reading tasks often require a huge amount of data to perform with similar accuracy to humans, yet it is also difficult to gain sufficient meaningful data. Automatic writing tasks have not been studied as extensively. In this paper, we teach machines to write like teaching a child by telling the machine the method for writing each character using L-attributed grammar. With the aid of the proposed TMTW (Teaching Machines To Write) interacting system, a human as a teacher only needs to provide the writing sequence of parts and control lines. The proposed system automatically perceives the relationships between control lines and parts, and constructs the grammars. Top-down derivation and the stroke generation method are applied to generate varying characters based on the learned grammars. For as long as a machine can write, it can be applied in robot control or training sample generation for automatic reading tasks. The MNIST and CASIA datasets are used to demonstrate the effectiveness of the proposed system on different languages. The machine written samples are used to train a network, which is evaluated on the MNIST test set. A test error rate of 1.23% is achieved using only approximately 20 grammars on average for each digit. Using the generated and handwritten samples together as a training set can reduce the test error rate to 0.61%. Similar experiments are conducted using the CASIA data set, and the results demonstrated that the proposed method is effective in generating characters with a complex structure. The source codes and grammars used in this paper have been made publicly available in https://github.com/step123456789/TMTW.