CompBio: Collaborative Research: Development of Effective Gene Selection Algorithms for Microarray Data Analysis
CompBio: Collaborative Research: Development of Effective Gene Selection Algorithms for Microarray Data Analysis
批准号:
0621889
负责人:
Haesun Park
金额:
$25.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2006
资助国家:
美国
项目状态:
已结题
起止时间:
2006-10-01 至 2010-09-30
中文摘要
随着人类基因组计划的成功,微阵列现在有可能在整个基因组范围内处理基因。一个典型的微阵列数据集包含大量的基因。对特定条件负责的重要基因数量的大幅减少,可以潜在地增加进一步生物学研究和关于特定基因作用的知识的可能性。任何能够提高我们在大量基因中识别重要基因的方法,以及通常有限的可用实验结果,都可能对我们对疾病和正常状态的理解产生重大影响,并最终对诊断、预后和药物设计产生重大影响。我们在这里提出的研究方法旨在提供有关基因作用的关键信息,其中我们的方法的关键组成部分是基于子空间的方法,该方法在许多模式识别任务中取得了巨大的成功,包括高效分类、聚类和快速搜索。开发有效的基于计算机的基因选择算法是必不可少的,因为由于问题的巨大复杂性,几乎不可能仅仅依靠生物测试。我们提出的研究的新颖和独特之处在于,我们试图找到一个数学上严谨的框架来模拟基因选择问题,并仔细考虑问题的生物学特征的重要性。利用我们在特征提取方面的知识和前人的研究成果,通过发现它们与特征选择之间的数学关系,设计出高效、有效的基因选择非参数方法。在我们提出的研究中,非负矩阵分解在建立特征提取和特征选择之间数学上严格的桥梁方面将发挥重要作用。在此过程中,我们还将探索基于交替最小二乘法和结构化总最小范数公式的估计缺失值的新方法,作为基因选择的预处理阶段。所有获得的结果、新算法和开发的软件,以及生成和编译的新数据集,将通过佐治亚理工学院和德克萨斯大学达拉斯分校现有的网络服务器,提供给研究界、教师、研究生和本科生。智力优势:这项研究将产生对计算微阵列分析有重大影响的方法。本研究开发的基因选择和缺失值估计方法,由于问题维数的初始化降低,大大降低了生物检测的复杂性,从而大大提高了对重要基因的详细研究。本研究开发的特征选择和特征提取算法将适用于许多其他需要高效处理高维空间数据集的问题,如文本处理、人脸识别、指纹分类、虹膜识别等。本研究设计的缺失值估计方法也可用于缺失数据的恢复,如协同过滤。更广泛的影响:这项研究将加强计算生物学和生物信息学的先进理论。所开发的技术在数据库管理、医学检查和诊断、生物化学选择和生物网络等方面也有潜在的应用。研究生参与这项研究将有许多未来的好处。学生的发现和研究经验将为他们在生物信息学中非常重要的当前研究领域的学术界,研究实验室和行业中富有成效的职业生涯做好准备。
英文摘要
With the success of the Human Genome Project, a microarray can now potentially handle the genes in an entire genome scale. A typical microarray data set involves a massive number of genes. A dramatic dimension reduction to a much smaller number of significant genes, responsible for specific conditions, can potentially increase the possibility of further biological study and knowledge regarding the roles of specific genes.Any methodology that can improve our recognition of significant genes among a large number of genes, and often a limited set of available experimental results, could have a significant impact on our understanding of diseased and normal states, and eventually on diagnosis, prognosis, and drug design. The method that we propose to investigate here is intended to provide critical information on the roles of genes where the key component of our approach is subspace-based methods, which have demonstrated great success in numerous pattern recognition tasks including efficient classification, clustering, and fast search.The development of effective computer-based algorithms for gene selection is indispensable since it is virtually impossible to rely solely on biological testing due to the enormous complexity of the problems. What is novel and unique in our proposed research is that we seek to find a mathematically rigorous framework that models gene selection problems, with careful consideration of the significance of the biological characteristics of the problem. Utilizing our knowledge and previous results on feature extraction, and by discovering their mathematical relationship to feature selection, efficient and effective nonparametric methods for gene selection will be designed. An important role will be played by the nonnegative matrix factorization in building a mathematically rigorous bridge between feature extraction and feature selection in our proposed research. In the process, we will also explore novel methods for estimating missing values as a preprocessing stage of gene selection based on the alternating least squares and the structured total least norm formulations. All results obtained, the new algorithms and software developed, as well as the new data sets generated and compiled will be made available to the research community, to teaching faculty, and to both graduate and undergraduate students, using existing Web servers at the Georgia Institute of Technology and University of Texas at Dallas.Intellectual Merit: This research will produce methods that will have a great impact on computational microarray analysis. The gene selection and missing value estimation methods developed in this research allow significant reduction in complexity of biological testing due to the initial reduction of the problem dimension, thus substantially improve detailed study of significant genes. The feature selection and feature extraction algorithms developed in this research will be applicable to many other problems where data sets in high dimensional spaces need to be handled efficiently and effectively, such as text processing, facial recognition, finger print classification, iris recognition. The missing value estimation methods designed in this research can also be utilized in recovering missing data such as in collaborative filtering.Broader Impact: The research will enhance advanced theory of computational biology and bioinformatics. The developed techniques will also have potential applications in database management, medical examination and diagnosis, bio-chemical selection, and biological networks. The graduate student involvement in this research will have numerous future benefits. The discovery and research experience of the students will prepare them for productive careers in academia, research labs, and industry in highly important current research areas in bioinformatics.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: OAC Core: Robust, Scalable, and Practical Low Rank Approximation
-
批准号:2106738
-
项目类别:Standard Grant
-
资助金额:$27.5万
-
财政年份:2021
-
负责人:Haesun Park
-
依托单位:
SI2-SSE: Collaborative Research: High Performance Low Rank Approximation for Scalable Data Analytics
-
批准号:1642410
-
项目类别:Standard Grant
-
资助金额:$33.23万
-
财政年份:2016
-
负责人:Haesun Park
-
依托单位:
CAREER: New Representations of Probability Distributions to Improve Machine Learning --- A Unified Kernel Embedding Framework for Distributions
-
批准号:1350983
-
项目类别:Continuing Grant
-
资助金额:$49.97万
-
财政年份:2014
-
负责人:Haesun Park
-
依托单位:
EAGER: Hierarchical Topic Modeling by Nonnegative Matrix Factorization for Interactive Multi-scale Analysis of Text Data
-
批准号:1348152
-
项目类别:Standard Grant
-
资助金额:$17.5万
-
财政年份:2013
-
负责人:Haesun Park
-
依托单位:
EAGER: Fast and Accurate Nonnegative Tensor Decompositions: Algorithms and Software
-
批准号:0956517
-
项目类别:Standard Grant
-
资助金额:$11.69万
-
财政年份:2009
-
负责人:Haesun Park
-
依托单位:
FODAVA-Lead: Dimension Reduction and Data Reduction: Foundations for Visualization
-
批准号:0808863
-
项目类别:Continuing Grant
-
资助金额:$300.0万
-
财政年份:2008
-
负责人:Haesun Park
-
依托单位:
MSPA-MCS: Collaborative Research: Fast Nonnegative Matrix Factorizations: Theory, Algorithms, and Applications
-
批准号:0732318
-
项目类别:Standard Grant
-
资助金额:$25.18万
-
财政年份:2007
-
负责人:Haesun Park
-
依托单位:
SGER: Effective Network Anomaly Detection Based on Adaptive Machine Learning
-
批准号:0715342
-
项目类别:Standard Grant
-
资助金额:$9.0万
-
财政年份:2007
-
负责人:Haesun Park
-
依托单位:
Collaborative Research: Greedy Approximations with Nonsubmodular Potential Functions
-
批准号:0728812
-
项目类别:Standard Grant
-
资助金额:$17.88万
-
财政年份:2007
-
负责人:Haesun Park
-
依托单位:
Special Meeting: Workshop on Future Direction in Numerical Algorithms and Optimization
-
批准号:0633793
-
项目类别:Standard Grant
-
资助金额:$0.0万
-
财政年份:2006
-
负责人:Haesun Park
-
依托单位:
Lower Dimensional Representation of Text Data for Efficient and Effective Information Retrieval
-
批准号:0549253
-
项目类别:Continuing Grant
-
资助金额:$0.0万
-
财政年份:2005
-
负责人:Haesun Park
-
依托单位:
ALGORITHMS: Collaborative Research: Development of Vector Space based Methods for Protein Structure Prediction
-
批准号:0549247
-
项目类别:Continuing Grant
-
资助金额:$0.0万
-
财政年份:2005
-
负责人:Haesun Park
-
依托单位:
Structure Preserving Reduced Rank Approximation: Theory, Algorithms and Software
-
批准号:9901992
-
项目类别:Standard Grant
-
资助金额:$16.1万
-
财政年份:1999
-
负责人:Haesun Park
-
依托单位:
Solution of Structured Total Least Norm and Parameter Estimation Problems
-
批准号:9509085
-
项目类别:Continuing Grant
-
资助金额:$21.86万
-
财政年份:1995
-
负责人:Haesun Park
-
依托单位:
Fast And Accurate Parallel Solutions for Recursive Least Squares Problems
-
批准号:9209726
-
项目类别:Standard Grant
-
资助金额:$10.44万
-
财政年份:1992
-
负责人:Haesun Park
-
依托单位:
海外基金