Towards the Integration of Natural Language and Eye Tracking Information for Predicting Comma Placement in Chinese Sentence
Towards the Integration of Natural Language and Eye Tracking Information for Predicting Comma Placement in Chinese Sentence
复制标题
DOI:
--
复制
发表时间:
2012
期刊:
影响因子:
--
通讯作者:
Chen Chen-Chen;Yoshinobu Kano;Akiko Aizawa
中科院分区:
文献类型:
--
作者:
Chen Chen-Chen;Yoshinobu Kano;Akiko Aizawa
This paper investigates a relatively underdeveloped but important subject in natural language processing (NLP) – prediction of punctuation marks. As a logographic language, Chinese provides a case of high contrast for alphabetic systems, because there are no word boundaries (spaces) between Chinese words. (Ren & Yang, 2010). It is sometimes difficult for Chinese learners to identify which characters compose certain words within a Chinese sentence. This characteristic raises an interesting question that how the segmentation cues such as spaces or commas inserted into text influence the course of Chinese reading. Bai et al. (2008) investigated spaced and unspaced Chinese reading but found no facilitation of spaces in word identification. As a visual mark in written language, a comma is widely used in Chinese writing to provide additional space and serve as a segmentation cue in Chinese text. More importantly, a comma not only has prosodic functions (Kerkhofs, et al. 2008) but can be placed at syntactic boundaries as well (Chafe, 1988). For example, a phrase or a clause in Chinese sentences is usually followed by a comma, although there is a great deal of flexibility in the use of commas. Ren & Yang (2010) showed that prosodic boundaries marked by commas facilitate sentence processing only when they cooperated with syntactic boundaries. In our present study, a Chinese comma placement predictor is expected to be created by integrating NLP and Eye Tracking technology. Lu & Ng (2010) proposed an approach built on top of dynamic conditional random fields (DCRF) framework, which jointly performs punctuation prediction together with sentence boundary and sentence type prediction on speech utterance without prosodic cues. Zhang, et al. (2009) presented a conditional random fields (CRF) based approach which automates ancient Chinese prose punctuation using the mutual information and the t-test difference as features. Guo & Wang (2010) tried combining sophisticated statistical techniques with linguistic analyses to facilitate generation of punctuation. However, how to detect a comma placement with prosodic functions remains a problem because it is more related with the human’s cognitive system but not only a linguistic problem. Eye tracking technology is considered to be a solution because it can show the intuitive reading process of people and detect hard-toread point when they are reading Chinese text without comma. After analyzing this difficulty, a comma distribution which more accords with the reader’s intuition is expected to be found. In this paper, we describe our comma predictor of modern Chinese with machine learning (ML) techniques by integrating several linguistic features so that the commas placed at syntactic boundaries can be predicted. We also briefly mention our plan of eye tracking experiments that examine whether eye tracking information can help find commas with prosodic function and then improve the accuracy of prediction. 2. Method