Are Readability Formulas Valid Tools for Assessing Survey Question Difficulty?

Are Readability Formulas Valid Tools for Assessing Survey Question Difficulty?
复制标题

DOI:
10.1177/0049124113513436
复制
发表时间:
2014-11-01
影响因子:
6.3
通讯作者:
Lenzner, Timo
Lenzner, Timo
中科院分区:
法学2区
文献类型:
--
作者:
Lenzner, Timo

文献摘要

被引文献

相似文献

可读性公式,如Flesch阅读容易公式,Flesch-Kincaid等级指数,Gunning Fog指数和Dale-Chall公式通常被认为是语言复杂性的客观度量。毫不奇怪,调查研究人员经常使用可读性分数作为问题难度的指标,并一再建议在问卷设计阶段应用这些公式,以确定有问题的项目,并协助调查设计人员修改有缺陷的问题。与此同时,这些公式在阅读研究者中受到了严厉的批评,特别是因为它们主要基于两个变量(单词长度/频率和句子长度),这可能不是语言困难的适当预测因素。本研究探讨上述四个可读性公式是否正确识别有问题的调查问题。可读性得分计算了71个问题对,其中每个问题包括一个有问题的(e。例如,在一个实施例中,句法复杂、模糊等)和问题的改进版本问题对来自两个来源:(1)现有的文献调查问卷设计和(2)Q-BANK数据库。分析表明,可读性公式往往有利于问题的改进版本。平均而言,公式在识别困难问题方面的成功率低于50%,并且各种公式之间的一致性差异很大。这种性能不佳的原因,以及在问卷设计和测试过程中使用的可读性公式的影响,进行了讨论。
Readability formulas, such as the Flesch Reading Ease formula, the Flesch-Kincaid Grade Level Index, the Gunning Fog Index, and the Dale-Chall formula are often considered to be objective measures of language complexity. Not surprisingly, survey researchers have frequently used readability scores as indicators of question difficulty and it has been repeatedly suggested that the formulas be applied during the questionnaire design phase, to identify problematic items and to assist survey designers in revising flawed questions. At the same time, the formulas have faced severe criticism among reading researchers, particularly because they are predominantly based on only two variables (word length/frequency and sentence length) that may not be appropriate predictors of language difficulty. The present study examines whether the four readability formulas named above correctly identify problematic survey questions. Readability scores were calculated for 71 question pairs, each of which included a problematic (e. g., syntactically complex, vague, etc.) and an improved version of the question. The question pairs came from two sources: (1) existing literature on questionnaire design and (2) the Q-BANK database. The analyses revealed that the readability formulas often favored the problematic over the improved version. On average, the success rate of the formulas in identifying the difficult questions was below 50 percent and agreement between the various formulas varied considerably. Reasons for this poor performance, as well as implications for the use of readability formulas during questionnaire design and testing, are discussed.