Moving beyond Kucera and Francis: A critical evaluation of current word frequency norms and the introduction of a new and improved word frequency measure for American English

Moving beyond Kucera and Francis: A critical evaluation of current word frequency norms and the introduction of a new and improved word frequency measure for American English
复制标题

DOI:
10.3758/brm.41.4.977
复制
发表时间:
2009-11-01
影响因子:
5.4
通讯作者:
New, Boris
New, Boris
中科院分区:
心理学2区
文献类型:
--
作者:
Brysbaert, Marc;New, Boris

文献摘要

被引文献

相似文献

词频是词加工和记忆研究中最重要的变量。然而,选择词频标准的主要标准一直是衡量标准的可用性,而不是其质量。因此,许多研究仍然基于旧的库塞拉和弗朗西斯频率规范。通过使用最近出版的超级研究的词汇决定时间,我们展示了这一衡量标准有多糟糕,以及必须做些什么来改进它。特别是,我们调查了语料库的大小,语料库所基于的语言语域,以及频率度量的定义。我们观察到,语料库大小对于小尺寸(取决于单词的使用频率)具有实际重要性,但对于超过1600万-3000万字的大小则不重要。至于语域,我们发现基于电视和电影字幕的频率要好于基于书面来源的频率,尤其是在心理语言学研究中使用的单音节和双音节单词。最后,我们发现在英语中词条频率并不优于词形频率,并且语境多样性的度量要好于基于原始出现频率的度量。后者的优越性在一定程度上是因为经常被用作名字的单词。在这些考虑的基础上汇编一个新的频率标准被证明比现有的标准(包括Kucera&Francis和CELEX)对字处理时间的预测要好得多。来自SUBTLEXUS语料库的新的SUBTL频率规范可从http://brm.psychonomic-journals.org/content/supplemental,以及根特大学和莱克斯克大学的网站免费获得,用于研究目的。
Word frequency is the most important variable in research on word processing and memory. Yet, the main criterion for selecting word frequency norms has been the availability of the measure, rather than its quality. As a result, much research is still based on the old Kucera and Francis frequency norms. By using the lexical decision times of recently published megastudies, we show how bad this measure is and what must be done to improve it. In particular, we investigated the size of the corpus, the language register on which the corpus is based, and the definition of the frequency measure. We observed that corpus size is of practical importance for small sizes (depending on the frequency of the word), but not for sizes above 16-30 million words. As for the language register, we found that frequencies based on television and film subtitles are better than frequencies based on written sources, certainly for the monosyllabic and bisyllabic words used in psycholinguistic research. Finally, we found that lemma frequencies are not superior to word form frequencies in English and that a measure of contextual diversity is better than a measure based on raw frequency of occurrence. Part of the superiority of the latter is due to the words that are frequently used as names. Assembling a new frequency norm on the basis of these considerations turned out to predict word processing times much better than did the existing norms (including Kucera & Francis and Celex). The new SUBTL frequency norms from the SUBTLEXUS corpus are freely available for research purposes from http://brm.psychonomic-journals.org/content/supplemental, as well as from the University of Ghent and Lexique Web sites.