Strategies for addressing collinearity in multivariate linguistic data

Strategies for addressing collinearity in multivariate linguistic data
复制标题

DOI:
10.1016/j.wocn.2018.09.004
复制
发表时间:
2018-11-01
影响因子:
1.9
通讯作者:
Baayen, R. Harald
Baayen, R. Harald
中科院分区:
人文科学1区
文献类型:
--
作者:
Tomaschek, Fabian;Hendrix, Peter;Baayen, R. Harald

文献摘要

被引文献

相似文献

当在回归建模中联合考虑多个相关预测因素时,估计系数可能会呈现出违反直觉且在理论上无法解释的值。我们综述了几种实现共线数据分析策略的统计方法:正则化回归(弹性网络)、监督成分广义线性回归和随机森林。举例说明了在德语语音语料库中具有大范围分段持续时间预测器的数据集的方法。结果大体上是一致的,但每种方法都有自己的优点和缺点。同时,它们为分析师提供了关于共线数据结构的略有不同但互补的观点。(C)2018年作者。爱思唯尔有限公司出版。
When multiple correlated predictors are considered jointly in regression modeling, estimated coefficients may assume counterintuitive and theoretically uninterpretable values. We survey several statistical methods that implement strategies for the analysis of collinear data: regression with regularization (the elastic net), supervised component generalized linear regression, and random forests. Methods are illustrated for a data set with a wide range of predictors for segment duration in a German speech corpus. Results broadly converge, but each method has its own strengths and weaknesses. Jointly, they provide the analyst with somewhat different but complementary perspectives on the structure of collinear data. (C) 2018 The Authors. Published by Elsevier Ltd.