课题基金 / 基金详情

CompBio: Collaborative Research: Development of Effective Gene Selection Algorithms for Microarray Data Analysis

CompBio: Collaborative Research: Development of Effective Gene Selection Algorithms for Microarray Data Analysis
CompBio:合作研究:开发用于微阵列数据分析的有效基因选择算法
批准号:
0621889
负责人:
Haesun Park
金额:
$25.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2006
资助国家:
美国
项目状态:
已结题
起止时间:
2006-10-01 至 2010-09-30

项目摘要

项目成果

Haesun Park的其他基金

相似基金

相关文献

中文摘要
翻译
随着人类基因组计划的成功,微阵列现在可以潜在地处理整个基因组规模的基因。典型的微阵列数据集涉及大量基因。一个显著的维度减少到一个更小的数量的显着基因,负责特定的条件,可以潜在地增加进一步的生物学研究和知识的可能性,关于特定基因的作用。任何方法,可以提高我们的识别显着基因在大量的基因,往往是有限的一组可用的实验结果,可能会对我们理解疾病和正常状态,并最终对诊断,预后和药物设计产生重大影响。我们在这里提出的研究方法旨在提供关于基因作用的关键信息,其中我们方法的关键组成部分是基于子空间的方法,这些方法在许多模式识别任务中取得了巨大成功,包括有效的分类,聚类,和快速搜索。有效的计算机的发展-基因选择的基于算法是必不可少的,因为由于问题的巨大复杂性,实际上不可能仅仅依靠生物测试。在我们提出的研究中,新颖和独特的是,我们试图找到一个数学上严格的框架来模拟基因选择问题,并仔细考虑问题的生物学特征的重要性。利用我们的知识和以前的结果特征提取,并通过发现它们的数学关系,特征选择,高效和有效的非参数方法基因选择将被设计。在我们提出的研究中,非负矩阵分解将在特征提取和特征选择之间建立一个数学上严格的桥梁中发挥重要作用。在这个过程中,我们还将探索新的方法估计缺失值作为基因选择的预处理阶段的交替最小二乘和结构化的总最小范数配方的基础上。所有获得的结果,新的算法和开发的软件,以及新的数据集生成和编译将提供给研究界,教师,研究生和本科生,使用现有的网络服务器在格鲁吉亚理工学院和德克萨斯大学达拉斯。智力优点:这项研究将产生的方法,将有很大的影响,计算微阵列分析。在本研究中开发的基因选择和缺失值估计方法允许显着降低生物测试的复杂性,由于初始的问题维度的减少,从而大大提高了重要基因的详细研究。本研究所开发的特征选择和特征提取算法将适用于许多其他需要高效处理高维空间数据集的问题,如文本处理、人脸识别、指纹分类、虹膜识别等。本研究设计的缺失值估计方法也可以用于恢复丢失的数据,如协同过滤。更广泛的影响:该研究将提高先进的计算生物学和生物信息学理论。所开发的技术也将在数据库管理、医学检查和诊断、生化选择和生物网络中具有潜在的应用。研究生参与这项研究将有许多未来的好处。学生的发现和研究经验将为他们在生物信息学当前非常重要的研究领域的学术界,研究实验室和工业界的生产性职业做好准备。
英文摘要
With the success of the Human Genome Project, a microarray can now potentially handle the genes in an entire genome scale. A typical microarray data set involves a massive number of genes. A dramatic dimension reduction to a much smaller number of significant genes, responsible for specific conditions, can potentially increase the possibility of further biological study and knowledge regarding the roles of specific genes.Any methodology that can improve our recognition of significant genes among a large number of genes, and often a limited set of available experimental results, could have a significant impact on our understanding of diseased and normal states, and eventually on diagnosis, prognosis, and drug design. The method that we propose to investigate here is intended to provide critical information on the roles of genes where the key component of our approach is subspace-based methods, which have demonstrated great success in numerous pattern recognition tasks including efficient classification, clustering, and fast search.The development of effective computer-based algorithms for gene selection is indispensable since it is virtually impossible to rely solely on biological testing due to the enormous complexity of the problems. What is novel and unique in our proposed research is that we seek to find a mathematically rigorous framework that models gene selection problems, with careful consideration of the significance of the biological characteristics of the problem. Utilizing our knowledge and previous results on feature extraction, and by discovering their mathematical relationship to feature selection, efficient and effective nonparametric methods for gene selection will be designed. An important role will be played by the nonnegative matrix factorization in building a mathematically rigorous bridge between feature extraction and feature selection in our proposed research. In the process, we will also explore novel methods for estimating missing values as a preprocessing stage of gene selection based on the alternating least squares and the structured total least norm formulations. All results obtained, the new algorithms and software developed, as well as the new data sets generated and compiled will be made available to the research community, to teaching faculty, and to both graduate and undergraduate students, using existing Web servers at the Georgia Institute of Technology and University of Texas at Dallas.Intellectual Merit: This research will produce methods that will have a great impact on computational microarray analysis. The gene selection and missing value estimation methods developed in this research allow significant reduction in complexity of biological testing due to the initial reduction of the problem dimension, thus substantially improve detailed study of significant genes. The feature selection and feature extraction algorithms developed in this research will be applicable to many other problems where data sets in high dimensional spaces need to be handled efficiently and effectively, such as text processing, facial recognition, finger print classification, iris recognition. The missing value estimation methods designed in this research can also be utilized in recovering missing data such as in collaborative filtering.Broader Impact: The research will enhance advanced theory of computational biology and bioinformatics. The developed techniques will also have potential applications in database management, medical examination and diagnosis, bio-chemical selection, and biological networks. The graduate student involvement in this research will have numerous future benefits. The discovery and research experience of the students will prepare them for productive careers in academia, research labs, and industry in highly important current research areas in bioinformatics.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: OAC Core: Robust, Scalable, and Practical Low Rank Approximation
  • 批准号:
    2106738
  • 项目类别:
    Standard Grant
  • 资助金额:
    $27.5万
  • 财政年份:
    2021
  • 负责人:
    Haesun Park
  • 依托单位:
SI2-SSE: Collaborative Research: High Performance Low Rank Approximation for Scalable Data Analytics
  • 批准号:
    1642410
  • 项目类别:
    Standard Grant
  • 资助金额:
    $33.23万
  • 财政年份:
    2016
  • 负责人:
    Haesun Park
  • 依托单位:
CAREER: New Representations of Probability Distributions to Improve Machine Learning --- A Unified Kernel Embedding Framework for Distributions
  • 批准号:
    1350983
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $49.97万
  • 财政年份:
    2014
  • 负责人:
    Haesun Park
  • 依托单位:
EAGER: Hierarchical Topic Modeling by Nonnegative Matrix Factorization for Interactive Multi-scale Analysis of Text Data
  • 批准号:
    1348152
  • 项目类别:
    Standard Grant
  • 资助金额:
    $17.5万
  • 财政年份:
    2013
  • 负责人:
    Haesun Park
  • 依托单位:
海外基金