SPR mega-benchmark shows surprisal tracks construction- but not item-level difficulty

SPR mega-benchmark shows surprisal tracks construction- but not item-level difficulty
复制标题

SPR 大型基准测试显示了令人惊讶的轨道构建 - 但不是项目级别的难度

DOI:
--
复制
发表时间:
2022
期刊:
影响因子:
--
通讯作者:
Tal Linzen
Tal Linzen
中科院分区:
--
文献类型:
--
作者:
Kuan;Suhas Arehalli;Mari Kugemoto;Christian Muxica;Grusha Prasad;Brian;Dillon;Tal Linzen

文献摘要

被引文献

相似文献

句法消歧困难的理论解释认为,词层面的不可预测性是决定花园小径句加工难度的唯一因素[1]。这意味着,语义分析应该预测跨不同结构和跨单个项目的处理难度。我们通过在我们介绍的大规模自定进度阅读(SPR)基准测试中观察结构级和项目级GP效应(GPE)来测试这一点:句法歧义处理(SAP)基准测试,该测试比标准阅读实验具有多个数量级的数据。我们专注于这个数据集中的一个构造子集:MV/RR(1a),NP/S(1b)和NP/Z(1c)。SAP基准使用标准的项目内析因设计来估计GPE。我们创造了24个句子,如(1)。每个参与者看到每个结构的4个模糊和4个明确的实例。这些句子与基准中的其他句子结构和30个不同的填充句子混合在一起。参与者在每个句子后回答一个理解问题;只有那些对填充物的准确率为80%或更高的人进行了分析(N=2000;在多产上招募)。每个项目有220-440个数据点,与以前的工作相比,产生了更精确的项目级GPE估计。首先,实证探索GP结构之间的相对难度,我们运行了三个贝叶斯最大结构的混合效应模型,每个模型用于消歧动词,第一个溢出词和第二个溢出词。在消歧动词中,NP/Z的GPE最大,其次是MV/RR,然后是NP/S(图1a)。对于所有构造,GPE在第一溢出位置处达到峰值,但与第一溢出位置相比,MV/RR构造中的峰值要高得多。
The surprisal account of syntactic disambiguation difficulty holds that word-level unpredictability is the sole determinant of processing difficulty in garden path (GP) sentences [1]. This entails that surprisal should predict processing difficulty both across different constructions and across individual items. We test this by looking at construction- and item-level GP effects (GPEs) in a large-scale self-paced-reading (SPR) benchmark we introduce: the Syntactic Ambiguity Processing (SAP) Benchmark, which has orders of magnitude more data than a standard reading experiment. We focus on a subset of the constructions in this dataset: MV/RR (1a), NP/S (1b), and NP/Z (1c)). The SAP benchmark uses a standard within-item factorial design to estimate GPEs. We created twenty-four sentence triplets as in (1). Each participant saw 4 ambiguous and 4 unambiguous instances of each construction. These sentences were intermixed with other sentence constructions from the benchmark and 30 diverse filler sentences. Participants answered a comprehension question following each sentence; only those whose accuracy on fillers was 80% or higher were analyzed (N=2000; recruited on Prolific). There were 220–440 datapoints per item, yielding more precise item-level estimates of GPEs compared to prior work. First, to empirically probe the relative difficulty among the GP constructions, we ran three Bayesian maximally-structured mixed-effect models, each for the disambiguating verb, the first spillover word and the second spillover word. At the disambiguating verb, NP/Z had the largest GPE, followed by MV/RR and then NP/S (Fig. 1a). For all constructions, GPEs peaked at the first spillover position, but the peak was much higher in the MV/RR construction compared to