Differential privacy-based evaporative cooling feature selection and classification with relief-F and random forests

Differential privacy-based evaporative cooling feature selection and classification with relief-F and random forests
复制标题

DOI:
10.1093/bioinformatics/btx298
复制
发表时间:
2017-09-15
期刊:
影响因子:
5.8
通讯作者:
McKinney, Brett A.
McKinney, Brett A.
中科院分区:
生物学3区
文献类型:
--
作者:
Le, Trang T.;Simmons, W. Kyle;McKinney, Brett A.

文献摘要

被引文献

相似文献

动机:从高维生物数据中以较低的预测误差将个体分类为疾病或临床类别是生物信息学中统计学习的重要挑战。特征选择可以提高分类精度,但必须小心地结合到交叉验证中,以避免过度匹配。最近,人们提出了基于差分隐私的特征选择方法,如差分私有随机森林和可重复使用的保持集。然而,对于生物信息学等领域,当特征数目远远大于观测数目p>n时,这些差异隐私方法容易被过度拟合。方法:我们引入了私有蒸发冷却算法,这是一种随机隐私保护机器学习算法,它使用Relipment-F进行特征选择,并使用随机森林进行隐私保护分类,同时也防止了过度拟合。我们将隐私保护门限机制与热力学Maxwell-Boltzmann分布联系起来,其中温度代表隐私门限。我们使用原子气体蒸发冷却的热统计物理概念进行反向逐步隐私保护特征选择。结果:在具有主效应和统计交互作用的模拟数据上,我们比较了三种隐私保护方法的保持和验证集的准确率:可重复使用的坚持、随机森林的重复坚持和私人蒸发冷却,其中私人蒸发冷却采用Relation-F特征选择和随机森林分类。在属性之间存在交互作用的模拟中,专用蒸发冷却提供了更高的分类精度,而不会基于独立的验证集进行过拟合。在没有相互作用的模拟中,随机森林和私人蒸发冷却的阈值提供了类似的精度。我们还将这些隐私方法应用于人类大脑静息状态功能磁共振数据,这些数据来自一项关于严重抑郁障碍的研究。可用性和实施:代码可通过http://insilico.utulsa.edu/software/privateEC.Contact:Brett-mckinney@utulsa.edu.补充信息:补充数据可在BioInformation Online上获得。
Motivation: Classification of individuals into disease or clinical categories from high-dimensional biological data with low prediction error is an important challenge of statistical learning in bioinformatics. Feature selection can improve classification accuracy but must be incorporated carefully into cross-validation to avoid overfitting. Recently, feature selection Methods based on differential privacy, such as differentially private random forests and reusable holdout sets, have been proposed. However, for domains such as bioinformatics, where the number of features is much larger than the number of observations p >> n, these differential privacy methods are susceptible to overfitting.Methods: We introduce private Evaporative Cooling, a stochastic privacy-preserving machine learning algorithm that uses Relief-F for feature selection and random forest for privacy preserving classification that also prevents overfitting. We relate the privacy-preserving threshold mechanism to a thermodynamic Maxwell-Boltzmann distribution, where the temperature represents the privacy threshold. We use the thermal statistical physics concept of Evaporative Cooling of atomic gases to perform backward stepwise privacy-preserving feature selection.Results: On simulated data with main effects and statistical interactions, we compare accuracies on holdout and validation sets for three privacy-preserving methods: the reusable holdout, reusable holdout with random forest, and private Evaporative Cooling, which uses Relief-F feature selection and random forest classification. In simulations where interactions exist between attributes, private Evaporative Cooling provides higher classification accuracy without overfitting based on an independent validation set. In simulations without interactions, thresholdout with random forest and private Evaporative Cooling give comparable accuracies. We also apply these privacy methods to human brain resting-state fMRI data from a study of major depressive disorder.Availability and implementation: Code available at http://insilico.utulsa.edu/software/privateEC.Contact: brett-mckinney@utulsa.eduSupplementary information: Supplementary data are available at Bioinformatics online.