A comparative study of variable selection methods in the context of developing psychiatric screening instruments.

A comparative study of variable selection methods in the context of developing psychiatric screening instruments.
复制标题

在开发精神筛查工具的背景下,可变选择方法的比较研究。

DOI:
10.1002/sim.5937
复制
发表时间:
2014-02-10
影响因子:
2
通讯作者:
Petkova, Eva
Petkova, Eva
中科院分区:
医学3区
文献类型:
--
作者:
Lu, Feihan;Petkova, Eva

文献摘要

参考文献

被引文献

相似文献

精神疾病筛查工具的开发涉及从现有的评估临床和行为表型的问卷中的项目池中选择项目。筛选工具应只包括几个项目,并在分类病例和非病例方面具有良好的准确性。在这种情况下可以使用变量/项目选择方法,例如最小绝对收缩和选择运算符(LASSO)、弹性网络、分类和回归树、随机森林和双样本t检验。与最常应用变量选择方法的情况不同(例如,超高维遗传或成像数据),精神病学数据通常具有较低的维度并且由以下因素表征:预测因子之间的相关性和可能的相互作用,重要变量的不可观察性(即,可用问卷未测量的真实变量)、预测因子中缺失值的数量和模式以及训练数据中病例的患病率。我们研究了这些因素如何影响几种变量选择方法的性能,并通过模拟比较它们的选择性能和预测错误率。我们的研究结果表明:(1)对于完全数据,LASSO和弹性网络在变量选择和未来数据预测方面优于其他方法;(2)对于某些类型的不完全数据,随机森林在插补中引入了偏差,导致变量重要性排序不正确,本文提出了结合随机森林插补和LASSO的插补-LASSO方法;这种方法抵消了随机森林中的偏差,并为缺失数据提供了一种简单而有效的项目选择方法。作为一个例子,我们应用的方法,从标准的自闭症诊断访谈修订版的项目。
The development of screening instruments for psychiatric disorders involves item selection from a pool of items in existing questionnaires assessing clinical and behavioral phenotypes. A screening instrument should consist of only a few items and have good accuracy in classifying cases and non-cases. Variable/item selection methods such as Least Absolute Shrinkage and Selection Operator (LASSO), Elastic Net, Classification and Regression Tree, Random Forest, and the two-sample t-test can be used in such context. Unlike situations where variable selection methods are most commonly applied (e.g., ultra high-dimensional genetic or imaging data), psychiatric data usually have lower dimensions and are characterized by the following factors: correlations and possible interactions among predictors, unobservability of important variables (i.e., true variables not measured by available questionnaires), amount and pattern of missing values in the predictors, and prevalence of cases in the training data. We investigate how these factors affect the performance of several variable selection methods and compare them with respect to selection performance and prediction error rate via simulations. Our results demonstrated that: (1) for complete data, LASSO and Elastic Net outperformed other methods with respect to variable selection and future data prediction, and (2) for certain types of incomplete data, Random Forest induced bias in imputation, leading to incorrect ranking of variable importance.We propose the Imputed-LASSO combining Random Forest imputation and LASSO; this approach offsets the bias in Random Forest and offers a simple yet efficient item selection approach for missing data. As an illustration, we apply the methods to items from the standard Autism Diagnostic Interview-Revised version.
DOI: 10.1038/tp.2012.10
发表时间: 2012-04-10
影响因子: 6.8
作者:
Wall DP;Kosmicki J;Deluca TF;Harstad E;Fusaro VA
通讯作者: Fusaro VA
DOI: 10.1038/sj.mp.4001034
发表时间: 2002-01-01
影响因子: 11
作者:
Knable, MB;Barci, BM;Torrey, EF
通讯作者: Torrey, EF
DOI: 10.1198/106186006x133933
发表时间: 2006-09-01
影响因子: 2.4
作者:
Hothorn, Torsten;Hornik, Kurt;Zeileis, Achim
通讯作者: Zeileis, Achim
DOI: 10.1214/07-aos520
发表时间: 2008-08-01
影响因子: 4.5
作者:
Zhang, Cun-Hui;Huang, Jian
通讯作者: Huang, Jian
DOI: 10.1176/appi.ajp.161.12.2322
发表时间: 2004-12-01
影响因子: 17.7
作者:
Fairburn, CG;Agras, WS;Stice, E
通讯作者: Stice, E