Knowing Factors or Factor Loadings, or Neither? Evaluating Estimators of Large Covariance Matrices with Noisy and Asynchronous Data

Knowing Factors or Factor Loadings, or Neither? Evaluating Estimators of Large Covariance Matrices with Noisy and Asynchronous Data
复制标题

DOI:
10.2139/ssrn.2920693
复制
发表时间:
2017-10
期刊:
Capital Markets: Market Microstructure eJournal
影响因子:
--
通讯作者:
Chaoxing Dai;Kun Lu;D. Xiu
Chaoxing Dai;Kun Lu;D. Xiu
中科院分区:
其他
文献类型:
--
作者:
Chaoxing Dai;Kun Lu;D. Xiu

文献摘要

被引文献

相似文献

我们调查估计因子模型为基础的大协方差(和精度)矩阵使用高频数据,这是异步的,并可能受到市场微观结构噪声的污染。我们的估计策略依赖于刷新时间的预平均方法来解决微观结构问题,同时分别使用三种不同规格的因子模型和各种阈值方法来对抗维数灾难。为了估计因子模型,如果因子是已知的,我们可以采用时间序列回归(TSR)来恢复载荷,或者使用横截面回归(CSR)来从已知的载荷中恢复因子,或者使用主成分分析(PCA),如果因子和它们的载荷都是未知的。我们比较了在这些情况下,使用联合填充和增加的维数渐近的收敛速度。为了评估所有30种估计策略组合对模型误定的鲁棒性和统计效率之间的经验权衡,我们分别对道琼斯30指数,标准普尔100指数和标准普尔500指数的样本外投资组合配置进行了赛马,并发现使用TSR或PCA的位置阈值占主导地位的基于预平均的策略,特别是基于子采样的替代方案。
We investigate estimators of factor-model-based large covariance (and precision) matrices using high-frequency data, which are asynchronous and potentially contaminated by the market microstructure noise. Our estimation strategies rely on the pre-averaging method with refresh time to solve the microstructure problems, while using three different specifications of factor models with a variety of thresholding methods, respectively, to battle the curse of dimensionality. To estimate a factor model, we either adopt the time-series regression (TSR) to recover loadings if factors are known, or use the cross-sectional regression (CSR) to recover factors from known loadings, or use the principal component analysis (PCA) if neither factors nor their loadings are assumed known. We compare the convergence rates in these scenarios using the joint in-fill and increasing dimensionality asymptotics. To evaluate the empirical trade-off between robustness to model misspecification and statistical efficiency among all 30 combinations of estimation strategies, we run a horse race on the out-of-sample portfolio allocation with Dow Jones 30, S&P 100, and S&P 500 index constituents, respectively, and find the pre-averaging-based strategy using TSR or PCA with location thresholding dominates, especially over the subsampling-based alternatives.