Dictionaries and distributions: Combining expert knowledge and large scale textual data content analysis

Dictionaries and distributions: Combining expert knowledge and large scale textual data content analysis
复制标题

DOI:
10.3758/s13428-017-0875-9
复制
发表时间:
2018-02-01
影响因子:
5.4
通讯作者:
Dehghani, Morteza
Dehghani, Morteza
中科院分区:
心理学2区
文献类型:
--
作者:
Garten, Justin;Hoover, Joe;Dehghani, Morteza

文献摘要

被引文献

相似文献

理论驱动的文本分析广泛使用了心理学概念词典,产生了广泛的重要结果。这些词典通常通过字数统计方法来应用,事实证明这种方法既简单又有效。在本文中,我们介绍了分布式字典表示(DDR),这是一种使用语义相似性而不是字数来应用心理词典的方法。这允许测量词典和从完整文档到单个单词的文本范围之间的相似性。我们展示了 DDR 如何使词典作者在不牺牲语言覆盖范围的情况下更加重视结构有效性。我们进一步证明了 DDR 在两个实际任务中的优势,最后对字典大小和任务性能之间的相互作用进行了广泛的研究。这些研究使我们能够研究 DDR 和字数统计方法作为应用概念词典的工具如何相互补充,以及每种方法的最佳应用场合。最后,我们提供了工具和资源的参考,以使这种方法可供广大心理受众使用和使用。
Theory-driven text analysis has made extensive use of psychological concept dictionaries, leading to a wide range of important results. These dictionaries have generally been applied through word count methods which have proven to be both simple and effective. In this paper, we introduce Distributed Dictionary Representations (DDR), a method that applies psychological dictionaries using semantic similarity rather than word counts. This allows for the measurement of the similarity between dictionaries and spans of text ranging from complete documents to individual words. We show how DDR enables dictionary authors to place greater emphasis on construct validity without sacrificing linguistic coverage. We further demonstrate the benefits of DDR on two real-world tasks and finally conduct an extensive study of the interaction between dictionary size and task performance. These studies allow us to examine how DDR and word count methods complement one another as tools for applying concept dictionaries and where each is best applied. Finally, we provide references to tools and resources to make this method both available and accessible to a broad psychological audience.