Vocabulary Learning Support System Based on Automatic Image Captioning Technology

Vocabulary Learning Support System Based on Automatic Image Captioning Technology
复制标题

DOI:
10.1007/978-3-030-21935-2_26
复制
发表时间:
2019-07
期刊:
--
影响因子:
--
通讯作者:
M. Hasnine;B. Flanagan;Gökhan Akçapınar;H. Ogata;Kousuke Mouri;Noriko Uosaki
M. Hasnine;B. Flanagan;Gökhan Akçapınar;H. Ogata;Kousuke Mouri;Noriko Uosaki
中科院分区:
其他
文献类型:
--
作者:
M. Hasnine;B. Flanagan;Gökhan Akçapınar;H. Ogata;Kousuke Mouri;Noriko Uosaki

文献摘要

相似文献

学习语境是词汇发展的重要组成部分,然而描述每个词汇的学习语境被认为是困难的。在人类大脑中,使用图片描述学习环境相对容易,因为图片可以快速描述文本注释无法做到的大量细节。因此,在一个非正式的语言学习系统中,图片可以用来克服语言学习者在描述学习环境时所面临的问题。本研究旨在开发一个支持系统,通过分析语言学习者拍摄的图片的视觉内容,自动生成和表示学习上下文。自动图像字幕是一种将计算机视觉和自然语言处理相结合的人工智能技术,用于分析学习者捕获的图像的视觉内容。一个名为Show and Tell的神经图像字幕生成器模型被训练用于图像到单词的生成,并描述图像的上下文。本研究的目标有三:第一,一种能理解图片内容并能自动生成学习情境的智能技术;第二,学习者可以通过使用一张图片来学习多个词汇,而不依赖于每个词汇的代表性图片,第三,学习者的先前词汇知识可以与新的学习词汇映射,使得在学习新词汇的同时回顾和回忆先前获得的词汇。
Learning context has evident to be an essential part in vocabulary development, however describing learning context for each vocabulary is considered to be difficult. In the human brain, it is relatively easy to describe learning contexts using pictures because pictures describe an immense amount of details at a quick glance that text annotations cannot do. Therefore, in an informal language learning system, pictures can be used to overcome the problems that language learners face in describing learning contexts. The present study aimed to develop a support system that generates and represents learning contexts automatically by analyzing the visual contents of the pictures captured by language learners. Automatic image captioning, a technology of artificial intelligence that connects computer vision and natural language processing is used for analyzing the visual contents of the learners’ captured images. A neural image caption generator model called Show and Tell is trained for image-to-word generation and to describe the context of an image. The three-fold objectives of this research are: First, an intelligent technology that can understand the contents of the picture and capable to generate learning contexts automatically; Second, a leaner can learn multiple vocabularies by using one picture without relying on a representative picture for each vocabulary, and Third, a learner’s prior vocabulary knowledge can be mapped with new learning vocabulary so that previously acquired vocabulary be reviewed and recalled while learning new vocabulary.