ITR: Optimal Support Set Selection in Data Analysis with Applications to Bioinformatics
ITR: Optimal Support Set Selection in Data Analysis with Applications to Bioinformatics
批准号:
0312953
负责人:
Peter Hammer
金额:
$0.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2003
资助国家:
美国
项目状态:
已结题
起止时间:
2003-08-01 至 2007-07-31
中文摘要
这项研究将开发系统的程序,利用计算机相关的发展和先进的组合优化技术,建立在以前成功的特别方法的基础上,优化数据分析中的特征选择,特别是生物信息学。从数据中提取知识是信息技术研究中的一个基本挑战。一种非常常见的知识提取问题是分析记录或观察的档案,以发现隐藏的结构关系。这种类型的问题出现在科学、技术、医学、管理的许多领域,以及无数其他活动领域。计算机和互联网的出现从根本上增加了数据分析的作用,不仅允许创建大型、有意义的数据集,而且使全球的研究人员都可以访问这些数据集。数据分析应用最突出的领域之一是系统生物学和生物信息学。与通常只关注单个分子的分子生物学研究不同,系统生物学同时关注数万甚至数十万个生物属性。此外,预计在不久的将来,数据集中包含的属性数量将大幅增加。新的技术使系统生物学的新的全球方法成为可能,这些技术允许同时测量大量属性和产生大型多参数数据集。这些生物数据集代表了生物信息学的领域。除了经典的统计方法外,还需要新的数据分析方法。其中,这些方法推动了全新研究领域的发展,包括机器学习、数据挖掘、神经网络、支持向量机等。在所有这些数据分析领域中,已知(有限)观测子集的知识被用来得出关于整个可能观测集合的结论。虽然在统计学的框架内以及主要基于组合学、逻辑和最优化的较新研究领域已经开发了强大的分析工具,但生物信息学领域(以及其他一些领域)的问题的规模导致了主要的计算困难,并提出了开发新的方法的挑战,其中系统的启发式程序被集成到求解算法中,以便智能地减小要解决的问题的规模。通过使用启发式算法以及组合学、逻辑学和最优化算法的特别组合,该项目的目标是在不显著损失结果模型的准确性的情况下实现问题规模的显著减少。该项目将引入新的概念来评估特征在数据分析问题中的作用和影响。这些概念结合了统计学、组合学、信息论、布尔函数理论、博弈论和投票理论。在算法方面,它们为消除冗余特征的系统启发式方法提供了可能性。该项目还将为特征对的比较评估引入新的概念,包括相似性和支配性特征。这些概念结合了统计学、偏序集理论和布尔代数中的元素,可以增加可用于消除不必要特征的工具的武器库。与以前的研究将属性集仅仅视为单个属性的集合不同,该项目提出了一种对属性的综合作用的研究。研究计划包括为一组属性制定一个局部最优标准,该标准应结合“支持集”所需的两个特征:它应允许建立准确的模型,同时应具有计算上可管理的规模。此外,该项目将引入新的综合“逻辑”属性,一方面允许压缩数据集,另一方面有可能找到明确可理解和实用的可用“逻辑”判别式,以区分正面和负面的意见。这些思想的一个非常重要的应用是提出的优化特征选择的算法框架。在生物研究和生物信息学领域,该项目引入:(I)“生物标记物组”的概念(可通过组合优化技术获得),(Ii)用于生物研究的算法假设生成器,以及(Iii)发现具有高度相似特征的新的观察类别的新方法。这项工作将产生一个公开可用的软件包,预计将促进数据计算分析和生物信息学的实质性研究。通过结合统计学、组合优化、信息论、布尔函数理论、对策、投票和偏序集理论,该项目有望吸引来自不同领域的研究人员进行合作研究。公开提供的生物信息学数据分析软件以及提出的假设生成系统将为促进生物学和生物信息学的研究提供主要工具。
英文摘要
This research will develop systematic procedures which take advantage of computer-related developments and advanced combinatorial optimization techniques, to build on previously successful ad-hoc methods for optimizing feature selection in data analysis, with special attention to bioinformatics. Knowledge extraction from data represents a fundamental challenge in information technology research. A very frequent type of knowledge extraction problem is that of analyzing archives of records or observations in order to discover hidden structural relationships. Problems of this type appear in numerous areas of science, technology, medicine, management, and in countless other areas of activity. The advent of the computer and of the Internet have radically increased the role of data analysis, by allowing not only the creation of large, meaningful datasets, but also by making them accessible to researchers all over the globe.One of the most prominent areas of applications of data analysis is in systems biology and bioinformatics. In contrast to molecular biology investigations, which typically focus on single molecules, systems biology pays attention to tens or even hundreds of thousands of biological attributes at the same time. Moreover, the number of attributes included in a dataset is predicted to increase dramatically in the very near future. The new global approach to systems biology has been enabled by new technologies that have allowed the simultaneous measurement of large numbers of attributes and the generation of large multiparameter datasets. These biological datasets represent the domain of bioinformatics. Beside the classic methods of statistics, new approaches to the analysis of data are required. Among others, these methodologies prompt the development of entirely new research areas, including e.g., machine learning, data mining, neural networks, support vector machines. In all these areas of data analysis, the knowledge of a known (finite) subset of observations is used to derive conclusions about the entire set of possible observations.While powerful analytic tools have been developed within the framework of statistics and of newer research areas based heavily on combinatorics, logic and optimization, the size of the problems in the area of bioinformatics (as well as in some other areas), leads to major computational difficulties, and raises the challenge of developing new approaches in which systematic heuristic procedures are integrated into solution algorithms, in order to intelligently reduce the size of the problems to be solved. By using ad-hoc combinations of heuristics and of combinatorics, logic, and optimization based algorithms, this project aims to achieve spectacular reductions of problem size without significant loss in the accuracy of the resulting models.This project will introduce new concepts for evaluating the role and the impact of features in data analysis problems. These concepts combine elements of statistics, combinatorics, information theory, the theory of Boolean functions, the theories of games and of voting. On the algorithmic side, they open possibilities for systematic heuristic approaches to the elimination of redundant features. The project will also introduce new concepts for the comparative evaluation of pairs of features, including those of similarity and domination. These concepts combine elements from statistics, the theory of partially ordered sets, and that of Boolean algebra, and can add to the arsenal of tools available for the elimination of unnecessary features. As opposed to previous studies which view the sets of attributes just as collections of individual attributes, the project proposes a study of the combined efforts of attributes. The research plan includes the development of a local optimality criterion for a set of attributes which should combine the two desired characteristics of a "support set": it should allow the construction of accurate models, and it should, at the same time, be of a computationally manageable size. In addition, the project will introduce new, synthetic "logical" attributes which allow, on the one hand, the compression of the dataset and, on the other hand, the possibility of finding clearly understandable and practical usable "logical" discriminants, which distinguish the positive observations from the negative ones. A very significant application of these ideas is the proposed algorithmic framework for optimizing feature selection. In the field of biological research and bioinformatics, the project introduces: (i) the concept of "groups of biomarkers" (obtainable through combinatorial optimization techniques), (ii) an algorithmic hypothesis generator for biological research, and (iii) a new approach for discovering new classes of observations with highly similar characteristics. A publicly available software package will result from the work and is expected to stimulate substantial research in the computational analysis of data, and in bioinformatics.By combining elements of statistics, combinatorial optimization, information theory, the theories of Boolean functions, games, voting, and partially ordered sets, the project is expected to attract researchers from a variety of areas to collaborative studies. The publicly available software for the analysis of bioinformatics data, as well as the hypothesis generation system proposed, will provide major tools for stimulating research in biology and bioinformatics.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Workshop on Discrete Optimization '99
-
批准号:9976754
-
项目类别:Standard Grant
-
资助金额:$1.5万
-
财政年份:1999
-
负责人:Peter Hammer
-
依托单位:
Pseudo-Boolean Functions: Representations and Optimization
-
批准号:9806389
-
项目类别:Standard Grant
-
资助金额:$17.5万
-
财政年份:1998
-
负责人:Peter Hammer
-
依托单位:
U.S.-Belgium Cooperative Research: Nonlinear 0-1 Optimization
-
批准号:9321811
-
项目类别:Standard Grant
-
资助金额:$1.5万
-
财政年份:1995
-
负责人:Peter Hammer
-
依托单位:
Mathematical Sciences Computing Research Environments
-
批准号:9406327
-
项目类别:Standard Grant
-
资助金额:$2.8万
-
财政年份:1994
-
负责人:Peter Hammer
-
依托单位:
Mathematical Sciences: Functions of Binary Variables
-
批准号:8906870
-
项目类别:Continuing Grant
-
资助金额:$13.05万
-
财政年份:1989
-
负责人:Peter Hammer
-
依托单位:
Nonlinear Binary Optimization
-
批准号:8503212
-
项目类别:Continuing Grant
-
资助金额:$16.11万
-
财政年份:1985
-
负责人:Peter Hammer
-
依托单位:
Mathematical Sciences: Structural and Algorithmic Aspects of Nonlinear Discrete Optimization
-
批准号:8305569
-
项目类别:Standard Grant
-
资助金额:$3.61万
-
财政年份:1984
-
负责人:Peter Hammer
-
依托单位:
海外基金