Machine learning prediction of cognition from functional connectivity: Are feature weights reliable?

Machine learning prediction of cognition from functional connectivity: Are feature weights reliable?
复制标题

DOI:
10.1016/j.neuroimage.2021.118648
复制
发表时间:
2021-10-22
期刊:
影响因子:
5.7
通讯作者:
Zalesky, Andrew
Zalesky, Andrew
中科院分区:
医学1区
文献类型:
--
作者:
Tian, Ye;Zalesky, Andrew

文献摘要

被引文献

相似文献

使用机器学习方法,可以从个人的功能性大脑连接中以适度的准确性预测认知表现。然而,到目前为止,预测模型对支持认知的神经生物学过程的见解有限。为此,特征选择和特征权重估计需要可靠,以确保能够可靠地识别具有高预测效用的重要连接和电路。我们全面研究了健康年轻人静息状态功能连接网络构建的各种认知表现预测模型的特征权重测试-重测可靠性(n=400)。尽管实现了适度的预测精度(r=0.2-0.4),但我们发现所有预测模型的特征权重可靠性通常较差(ICC< 0.3),并且明显低于明显的生物属性(如性别)的预测模型(ICC接近0.5)。较大的样本量(n=800)、Haufe变换、非稀疏特征选择/正则化和较小的特征空间略微提高了可靠性(ICC< 0.4)。我们阐明了特征权重可靠性和预测精度之间的权衡,并发现单变量统计比预测模型的特征权重略微更可靠。最后,我们证明了在交叉验证折叠之间测量特征权重的一致性提供了对特征权重可靠性的夸大估计。因此,如果可能的话,我们建议在样本外估计可靠性。我们认为,重新平衡焦点从预测准确性到模型可靠性可能有助于机器学习方法对认知的机制理解。
Cognitive performance can be predicted from an individual's functional brain connectivity with modest accuracy using machine learning approaches. As yet, however, predictive models have arguably yielded limited insight into the neurobiological processes supporting cognition. To do so, feature selection and feature weight estimation need to be reliable to ensure that important connections and circuits with high predictive utility can be reliably identified. We comprehensively investigate feature weight test-retest reliability for various predictive models of cognitive performance built from resting-state functional connectivity networks in healthy young adults (n=400). Despite achieving modest prediction accuracies (r=0.2-0.4), we find that feature weight reliability is generally poor for all predictive models (ICC< 0.3), and significantly poorer than predictive models for overt biological attributes such as sex (ICC approximate to 0.5). Larger sample sizes (n=800), the Haufe transformation, non-sparse feature selection/regularization and smaller feature spaces marginally improve reliability (ICC< 0.4). We elucidate a tradeoff between feature weight reliability and prediction accuracy and find that univariate statistics are marginally more reliable than feature weights from predictive models. Finally, we show that measuring agreement in feature weights between cross-validation folds provides inflated estimates of feature weight reliability. We thus recommend for reliability to be estimated out-of-sample, if possible. We argue that rebalancing focus from prediction accuracy to model reliability may facilitate mechanistic understanding of cognition with machine learning approaches.