Reading the written language environment: Learning orthographic structure from statistical regularities

Reading the written language environment: Learning orthographic structure from statistical regularities
复制标题

DOI:
10.1016/j.jml.2020.104148
复制
发表时间:
2020-10
影响因子:
4.3
通讯作者:
Teresa M. Schubert;T. Cohen;S. Fischer-Baum
Teresa M. Schubert;T. Cohen;S. Fischer-Baum
中科院分区:
心理学2区
文献类型:
--
作者:
Teresa M. Schubert;T. Cohen;S. Fischer-Baum

文献摘要

被引文献

相似文献

环境中的统计数据影响跨领域的认知。在语义学中,分布方法认为词与词之间的相似性可以从它们出现的上下文的相似性中得到。在这里,我们研究了书面文本中的重复如何影响读者对正字法的认识:字符之间的相似性可以从书面环境中学习吗?采用分布语义学的方法,我们在一个大型文本语料库中建立了字母数字字符之间的上下文相似度模型。我们发现适度的相关性模型衍生的相似性来自行为实验。除此之外,来自神经嵌入模型的模型衍生相似性捕获了拼写知识的关键方面,如大小写,字母身份和辅音-元音状态。我们的结论是,文本环境包含与读者相关的信息,统计学习是一个很有前途的方式来获取这些信息。更广泛地说,我们的研究结果意味着,统计学意义相关的不仅在单词语义的水平,但也个别书面字符。
Statistical regularities in the environment impact cognition across domains. In semantics, distributional approaches posit that similarity between words can be derived from regularities of the contexts in which they appear. Here, we study how regularities in written text impact readers’ knowledge about orthography: Can similarity between characters be learned from the written environment? Adapting methods from distributional semantics, we model the contextual similarity among alphanumeric characters in a large text corpus. We find modest correlations between model-derived similarities with similarity derived from a behavioral experiment. Beyond this result, model-derived similarity from neural embedding models captures key aspects of orthographic knowledge, like case, letter identity and consonant–vowel status. We conclude that the text environment contains regularities that are relevant to readers and that statistical learning is a promising way for this information to be acquired. More broadly, our results imply that statistical regularities are relevant not only at the level of word semantics but also individual written characters.