RANK DISCRIMINANTS FOR PREDICTING PHENOTYPES FROM RNA EXPRESSION

RANK DISCRIMINANTS FOR PREDICTING PHENOTYPES FROM RNA EXPRESSION
复制标题

DOI:
10.1214/14-aoas738
复制
发表时间:
2014-09-01
影响因子:
1.8
通讯作者:
Geman, Donald
Geman, Donald
中科院分区:
数学4区
文献类型:
--
作者:
Afsari, Bahman;Braga-Neto, Ulisses M.;Geman, Donald

文献摘要

被引文献

相似文献

在计算生物学中,分析大规模生物分子数据的统计方法是很常见的。一个值得注意的例子是根据基因表达数据进行表型预测,例如,检测人类癌症、区分亚型和预测临床结果。尽管如此,临床应用仍然很少。一个原因是,标准统计学习产生的决策规则的复杂性阻碍了生物学理解,特别是任何机械性解释。在这里,我们只利用几个基因之间的表达顺序来探索二进制分类的决策规则;然后,基本的构建块是两个基因的表达比较。最简单的例子,只是一个比较,是TSP分类器,它已经出现在各种与癌症相关的发现研究中。基于多重比较的决策规则可以更好地适应类的异质性,从而提高准确性,并可能提供与生物机制的联系。我们考虑了一个用于设计判别函数的通用框架(“上下文中的排名”),包括对支持(“上下文”)中的基因的数目和身份的数据驱动选择。然后我们专门举了两个例子:在几对基因中投票,并比较两组基因的中位数表达。全面的实验评估了相对于其他更复杂的方法的准确性,并强化了早期的观察结果,即简单的分类器是有竞争力的。
Statistical methods for analyzing large-scale biomolecular data are commonplace in computational biology. A notable example is phenotype prediction from gene expression data, for instance, detecting human cancers, differentiating subtypes and predicting clinical outcomes. Still, clinical applications remain scarce. One reason is that the complexity of the decision rules that emerge from standard statistical learning impedes biological understanding, in particular, any mechanistic interpretation. Here we explore decision rules for binary classification utilizing only the ordering of expression among several genes; the basic building blocks are then two-gene expression comparisons. The simplest example, just one comparison, is the TSP classifier, which has appeared in a variety of cancer-related discovery studies. Decision rules based on multiple comparisons can better accommodate class heterogeneity, and thereby increase accuracy, and might provide a link with biological mechanism. We consider a general framework ("rank-in-context") for designing discriminant functions, including a data-driven selection of the number and identity of the genes in the support ("context"). We then specialize to two examples: voting among several pairs and comparing the median expression in two groups of genes. Comprehensive experiments assess accuracy relative to other, more complex, methods, and reinforce earlier observations that simple classifiers are competitive.