Framework for kernel regularization with application to protein clustering

Framework for kernel regularization with application to protein clustering
复制标题

DOI:
10.1073/pnas.0505411102
复制
发表时间:
2005-08-30
影响因子:
11.1
通讯作者:
Wahba, G
Wahba, G
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Lu, F;Keles, S;Wahba, G

文献摘要

被引文献

相似文献

我们开发和应用一个以前未描述的框架,该框架旨在从可能粗糙,嘈杂,不完整,不一致的对象对之间的相异性信息,可在各种情况下,以正定核矩阵的形式提取信息。任何正定核定义了一组一致的距离,拟合核提供了一组欧几里得空间中的坐标,试图在控制核的复杂性的同时考虑可用的信息。所得到的坐标集非常适合于可视化,并作为分类和聚类算法的输入。该框架制定的一类优化问题,可以有效地解决使用现代凸锥规划软件。该方法的功率示出在基于一级序列数据的蛋白质聚类的上下文中。珠蛋白家族的蛋白质的应用程序导致在一个容易可视化的3D序列空间的珠蛋白,其中几个亚家族和亚组与文献一致,很容易识别。
We develop and apply a previously undescribed framework that is designed to extract information in the form of a positive definite kernel matrix from possibly crude, noisy, incomplete, inconsistent dissimilarity information between pairs of objects, obtainable in a variety of contexts. Any positive definite kernel defines a consistent set of distances, and the fitted kernel provides a set of coordinates in Euclidean space that attempts to respect the information available while controlling for complexity of the kernel. The resulting set of coordinates is highly appropriate for visualization and as input to classification and clustering algorithms. The framework is formulated in terms of a class of optimization problems that can be solved efficiently by using modern convex cone programming software. The power of the method is illustrated in the context of protein clustering based on primary sequence data. An application to the globin family of proteins resulted in a readily visualizable 3D sequence space of globins, where several subfamilies and subgroupings consistent with the literature were easily identifiable.