Scoring best-worst data in unbalanced many-item designs, with applications to crowdsourcing semantic judgments

Scoring best-worst data in unbalanced many-item designs, with applications to crowdsourcing semantic judgments
复制标题

DOI:
10.3758/s13428-017-0898-2
复制
发表时间:
2018-04-01
影响因子:
5.4
通讯作者:
Hollis, Geoff
Hollis, Geoff
中科院分区:
心理学2区
文献类型:
--
作者:
Hollis, Geoff

文献摘要

被引文献

相似文献

最佳-最差评分是一种判断格式,其中参与者被呈现一组项目,并且必须选择该组中的上级和下级项目。最佳-最差缩放在每次判断中产生大量信息,因为每次判断都允许对所有未判断项目的排名值进行推断。这种最佳-最差缩放的特性使其成为心理学和自然语言处理研究中一种很有前途的判断格式,这些研究涉及估计数万个单词的语义属性。各种不同的评分算法已经在以前的文献中设计的最好的最坏的缩放。然而,由于计算效率的问题,这些评分算法不能有效地应用于需要对数千个项目进行评分的情况。这里提出了新的算法,用于将响应从最佳最差缩放到数千个项目的项目分数(多项目评分问题)。这些评分算法通过模拟和实证实验进行验证,并确定了与噪声,真实值的基本分布和试验设计相关的考虑因素,这些因素可能会影响衍生项目评分的相对质量。新引入的评分算法始终优于以往文献中使用的评分算法对多项最佳-最差数据进行评分。
Best-worst scaling is a judgment format in which participants are presented with a set of items and have to choose the superior and inferior items in the set. Best-worst scaling generates a large quantity of information per judgment because each judgment allows for inferences about the rank value of all unjudged items. This property of best-worst scaling makes it a promising judgment format for research in psychology and natural language processing concerned with estimating the semantic properties of tens of thousands of words. A variety of different scoring algorithms have been devised in the previous literature on best-worst scaling. However, due to problems of computational efficiency, these scoring algorithms cannot be applied efficiently to cases in which thousands of items need to be scored. New algorithms are presented here for converting responses from best-worst scaling into item scores for thousands of items (many-item scoring problems). These scoring algorithms are validated through simulation and empirical experiments, and considerations related to noise, the underlying distribution of true values, and trial design are identified that can affect the relative quality of the derived item scores. The newly introduced scoring algorithms consistently outperformed scoring algorithms used in the previous literature on scoring many-item best-worst data.