Sequence-Based Classification Using Discriminatory Motif Feature Selection

Sequence-Based Classification Using Discriminatory Motif Feature Selection
复制标题

DOI:
10.1371/journal.pone.0027382
复制
发表时间:
2011-11-10
期刊:
影响因子:
3.7
通讯作者:
Segal, Mark R.
Segal, Mark R.
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Xiong, Hao;Capurso, Daniel;Segal, Mark R.

文献摘要

被引文献

相似文献

大多数现有的基于序列的分类方法使用穷举特征生成,例如,采用所有的k-mer模式。这种(枚举)方法背后的动机是最大限度地减少忽略重要特征的可能性。然而,这一战略也存在缺陷。首先,实际约束将穷举特征生成的范围限制到长度为k的模式,不考虑预测器。其次,这样生成的特征表现出很强的依赖性,这可能使对派生分类规则的理解变得复杂。第三,也是最重要的一点,大量不相关的特征被创造出来。这些担忧可能会影响预测和解释。虽然提出了补救办法,但这些办法往往针对具体问题,不能广泛适用。在这里,我们开发了一个普遍适用的方法,以及随之而来的软件管道,这是基于歧视性的基序发现。除了传统的训练和验证分区之外,我们的框架还需要第三层的数据分区,即发现分区。在发现分区中的序列和相关类别标签上使用歧视性基序查找器以产生(小)特征集。然后,这些特征被用作训练分区中的分类器的输入。最后,在验证分区上进行性能评估。我们的方法的重要属性是其模块性(可以部署任何歧视性基序发现器和任何分类器)和其通用性(可以容纳所有数据,包括未对齐和/或长度不等的序列)。我们说明了我们的方法上的两个核小体占用数据集和蛋白质溶解度数据集,以前使用枚举功能生成分析。我们的方法实现了出色的性能结果,有和没有优化的分类器调整参数。实现该方法的Python管道可在http://www.epibiostat.ucsf.edu/biostat/sen/dmfs/上获得。
Most existing methods for sequence-based classification use exhaustive feature generation, employing, for example, all k-mer patterns. The motivation behind such (enumerative) approaches is to minimize the potential for overlooking important features. However, there are shortcomings to this strategy. First, practical constraints limit the scope of exhaustive feature generation to patterns of length k) predictors are not considered. Second, features so generated exhibit strong dependencies, which can complicate understanding of derived classification rules. Third, and most importantly, numerous irrelevant features are created. These concerns can compromise prediction and interpretation. While remedies have been proposed, they tend to be problem-specific and not broadly applicable. Here, we develop a generally applicable methodology, and an attendant software pipeline, that is predicated on discriminatory motif finding. In addition to the traditional training and validation partitions, our framework entails a third level of data partitioning, a discovery partition. A discriminatory motif finder is used on sequences and associated class labels in the discovery partition to yield a (small) set of features. These features are then used as inputs to a classifier in the training partition. Finally, performance assessment occurs on the validation partition. Important attributes of our approach are its modularity (any discriminatory motif finder and any classifier can be deployed) and its universality (all data, including sequences that are unaligned and/or of unequal length, can be accommodated). We illustrate our approach on two nucleosome occupancy datasets and a protein solubility dataset, previously analyzed using enumerative feature generation. Our method achieves excellent performance results, with and without optimization of classifier tuning parameters. A Python pipeline implementing the approach is available at http://www.epibiostat.ucsf.edu/biostat/sen/dmfs/.