Testing the Robustness of Online Word Segmentation: Effects of Linguistic Diversity and Phonetic Variation

Testing the Robustness of Online Word Segmentation: Effects of Linguistic Diversity and Phonetic Variation
复制标题

测试在线分词的稳健性:语言多样性和语音变异的影响

DOI:
--
复制
发表时间:
2011
期刊:
CMCL@ACL
影响因子:
--
通讯作者:
Emmanuel Dupoux
Emmanuel Dupoux
中科院分区:
--
文献类型:
--
作者:
Luc Boruta;S. Peperkamp;Benoît Crabbé;Emmanuel Dupoux

文献摘要

被引文献

相似文献

通常使用语音转录语料库评估获得单词细分的模型。因此,他们隐含地假设孩子在学会从语音中提取单词时如何撤消语音变化。此外,尽管语言获取模型应在语言上相似,但评估通常仅限于英语样本。我们使用以儿童为指导的英语,法语和日语语料库,评估了鉴于尚未减少语音变化的输入的最先进统计模型的性能。为此,我们测量了分段稳健性,跨不同级别的分段变化,模拟系统的同性音量变化或音素识别中的误差。我们表明,这些模型不能抗拒这种变化的增加,也不会推广到类型不同的语言。从早期语言获取的角度来看,结果加强了假设,根据词典构建词典之前在很大程度上获得语音知识的假设。
Models of the acquisition of word segmentation are typically evaluated using phonemically transcribed corpora. Accordingly, they implicitly assume that children know how to undo phonetic variation when they learn to extract words from speech. Moreover, whereas models of language acquisition should perform similarly across languages, evaluation is often limited to English samples. Using child-directed corpora of English, French and Japanese, we evaluate the performance of state-of-the-art statistical models given inputs where phonetic variation has not been reduced. To do so, we measure segmentation robustness across different levels of segmental variation, simulating systematic allophonic variation or errors in phoneme recognition. We show that these models do not resist an increase in such variations and do not generalize to typologically different languages. From the perspective of early language acquisition, the results strengthen the hypothesis according to which phonological knowledge is acquired in large part before the construction of a lexicon.