Towards the Integration of Natural Language and Eye Tracking Information for Predicting Comma Placement in Chinese Sentence

Towards the Integration of Natural Language and Eye Tracking Information for Predicting Comma Placement in Chinese Sentence
复制标题

DOI:
--
复制
发表时间:
2012
期刊:
--
影响因子:
--
通讯作者:
Chen Chen-Chen;Yoshinobu Kano;Akiko Aizawa
Chen Chen-Chen;Yoshinobu Kano;Akiko Aizawa
中科院分区:
其他
文献类型:
--
作者:
Chen Chen-Chen;Yoshinobu Kano;Akiko Aizawa

文献摘要

相似文献

本文研究了自然语言处理中一个相对不发达但很重要的课题--标点符号的预测。作为一种标志语言,汉语为字母系统提供了一个高对比度的案例,因为中文单词之间没有词的边界(空格)。(任和杨,2010)。对于汉语学习者来说,有时很难识别出一个汉语句子中的某些单词是由哪些汉字组成的。这一特点提出了一个有趣的问题,即文本中插入空格或逗号等切分线索如何影响汉语阅读过程。白等人。(2008)研究了有间隔和无间隔的中文阅读,但没有发现空格在单词识别中的促进作用。逗号作为书面语中的一种视觉符号,在汉语写作中被广泛使用,以提供额外的空间,并在汉语文本中作为切分线索。更重要的是,逗号不仅具有韵律功能(Kerkhoff等人)。2008),但也可以置于句法边界(Chafe,1988)。例如,汉语句子中的短语或从句通常后面都有逗号,尽管逗号的使用有很大的灵活性。任和杨(2010)指出,只有当韵律边界与句法边界相配合时,逗号才能促进句子的加工。在我们目前的研究中,我们希望通过整合自然语言处理和眼动跟踪技术来创建一个中文逗号位置预测器。Lu&Ng(2010)提出了一种建立在动态条件随机场(DCRF)框架之上的方法,该方法结合句子边界和句子类型预测对无韵律线索的语音进行标点符号预测和句子类型预测。Zhang,et al.(2009)提出了一种基于条件随机场(CRF)的古文标点自动标注方法,该方法以互信息和t检验差为特征。Guo&Wang(2010)尝试将复杂的统计技术与语言分析相结合,以促进标点符号的生成。然而,如何检测具有韵律功能的逗号位置仍然是一个问题,因为它更多地与人类的认知系统有关,而不仅仅是语言问题。眼动跟踪技术可以直观地显示人的阅读过程,在阅读无逗号的中文文本时发现难读点,因此被认为是一种解决方案。在分析了这一难点之后,期望找到一种更符合读者直觉的逗号分布。在本文中,我们使用机器学习(ML)技术描述了我们的现代汉语逗号预测器,它综合了几个语言特征,从而能够预测位于句法边界的逗号。我们还简要介绍了我们的眼睛跟踪实验计划,该实验检查眼睛跟踪信息是否可以帮助找到具有韵律功能的逗号,从而提高预测的准确性。2.方法
This paper investigates a relatively underdeveloped but important subject in natural language processing (NLP) – prediction of punctuation marks. As a logographic language, Chinese provides a case of high contrast for alphabetic systems, because there are no word boundaries (spaces) between Chinese words. (Ren & Yang, 2010). It is sometimes difficult for Chinese learners to identify which characters compose certain words within a Chinese sentence. This characteristic raises an interesting question that how the segmentation cues such as spaces or commas inserted into text influence the course of Chinese reading. Bai et al. (2008) investigated spaced and unspaced Chinese reading but found no facilitation of spaces in word identification. As a visual mark in written language, a comma is widely used in Chinese writing to provide additional space and serve as a segmentation cue in Chinese text. More importantly, a comma not only has prosodic functions (Kerkhofs, et al. 2008) but can be placed at syntactic boundaries as well (Chafe, 1988). For example, a phrase or a clause in Chinese sentences is usually followed by a comma, although there is a great deal of flexibility in the use of commas. Ren & Yang (2010) showed that prosodic boundaries marked by commas facilitate sentence processing only when they cooperated with syntactic boundaries. In our present study, a Chinese comma placement predictor is expected to be created by integrating NLP and Eye Tracking technology. Lu & Ng (2010) proposed an approach built on top of dynamic conditional random fields (DCRF) framework, which jointly performs punctuation prediction together with sentence boundary and sentence type prediction on speech utterance without prosodic cues. Zhang, et al. (2009) presented a conditional random fields (CRF) based approach which automates ancient Chinese prose punctuation using the mutual information and the t-test difference as features. Guo & Wang (2010) tried combining sophisticated statistical techniques with linguistic analyses to facilitate generation of punctuation. However, how to detect a comma placement with prosodic functions remains a problem because it is more related with the human’s cognitive system but not only a linguistic problem. Eye tracking technology is considered to be a solution because it can show the intuitive reading process of people and detect hard-toread point when they are reading Chinese text without comma. After analyzing this difficulty, a comma distribution which more accords with the reader’s intuition is expected to be found. In this paper, we describe our comma predictor of modern Chinese with machine learning (ML) techniques by integrating several linguistic features so that the commas placed at syntactic boundaries can be predicted. We also briefly mention our plan of eye tracking experiments that examine whether eye tracking information can help find commas with prosodic function and then improve the accuracy of prediction. 2. Method