Hierarchical Feature Selection Incorporating Known and Novel Biological Information: Identifying Genomic Features Related to Prostate Cancer Recurrence.

Hierarchical Feature Selection Incorporating Known and Novel Biological Information: Identifying Genomic Features Related to Prostate Cancer Recurrence.
复制标题

DOI:
10.1080/01621459.2016.1164051
复制
发表时间:
2016
影响因子:
3.7
通讯作者:
Long Q
Long Q
中科院分区:
数学1区
文献类型:
--
作者:
Zhao Y;Chung M;Johnson BA;Moreno CS;Long Q

文献摘要

被引文献

相似文献

我们的工作是由一项前列腺癌研究的动机,旨在确定mRNA和miRNA生物标志物,预测前列腺切除术后癌症复发。在文献中已经显示,结合关于途径成员和生物标志物之间的相互作用的已知生物信息改善了与疾病风险相关的高维生物标志物的特征选择。生物信息通常由图或网络表示,其中生物标志物由节点表示,它们之间的相互作用由边表示;然而,生物信息通常不完全已知。例如,microRNA(miRNAs)在调控基因表达中的作用尚未完全理解,并且miRNA调控网络尚未完全建立,在这种情况下,需要新的策略进行特征选择。为此,我们将未知的生物信息视为缺失数据(即,图中的缺失边),这不同于通常遇到的变量值缺失的缺失数据问题。本文提出了基于观测数据的未知生物信息估算的新概念,并将估算的信息定义为新的生物信息。此外,我们提出了一个分层组惩罚,以鼓励稀疏性和特征选择在两个途径水平和内途径水平,这与插补步骤相结合,允许纳入已知的和新的生物信息。虽然它是适用于一般的回归设置,我们开发和研究所提出的方法的背景下,半参数加速失效时间模型的动机,我们的数据示例。数据应用和仿真研究表明,新的生物信息的结合提高了风险预测和特征选择的性能,所提出的惩罚优于现有的几种惩罚的扩展。
Our work is motivated by a prostate cancer study aimed at identifying mRNA and miRNA biomarkers that are predictive of cancer recurrence after prostatectomy. It has been shown in the literature that incorporating known biological information on pathway memberships and interactions among biomarkers improves feature selection of high-dimensional biomarkers in relation to disease risk. Biological information is often represented by graphs or networks, in which biomarkers are represented by nodes and interactions among them are represented by edges; however, biological information is often not fully known. For example, the role of microRNAs (miRNAs) in regulating gene expression is not fully understood and the miRNA regulatory network is not fully established, in which case new strategies are needed for feature selection. To this end, we treat unknown biological information as missing data (i.e., missing edges in graphs), different from commonly encountered missing data problems where variable values are missing. We propose a new concept of imputing unknown biological information based on observed data and define the imputed information as the novel biological information. In addition, we propose a hierarchical group penalty to encourage sparsity and feature selection at both the pathway level and the within-pathway level, which, combined with the imputation step, allows for incorporation of known and novel biological information. While it is applicable to general regression settings, we develop and investigate the proposed approach in the context of semiparametric accelerated failure time models motivated by our data example. Data application and simulation studies show that incorporation of novel biological information improves performance in risk prediction and feature selection and the proposed penalty outperforms the extensions of several existing penalties.