ITR Collaborative Research: Combinatorial Algorithms for Biological Data Clustering
ITR Collaborative Research: Combinatorial Algorithms for Biological Data Clustering
批准号:
0325386
负责人:
Ying Xu
金额:
$129.5万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2003
资助国家:
美国
项目状态:
已结题
起止时间:
2003-09-15 至 2004-05-31
中文摘要
项目摘要人类基因组计划打开了生物数据的闸门,导致大量序列、结构、表达和相互作用数据的生成,其速度远远超出了我们目前分析和解释这些数据的能力。迫切需要新的想法和方法来大大提高生物数据分析的能力。数据聚类是挖掘大量生物数据的基础。该项目的目标是(a)开发一个高效且通用的生物数据聚类框架,适用于一大类生物数据分析问题; (b) 通过应用于四个具有挑战性的生物数据分析问题,证明该框架作为通用聚类工具的有效性; (c) 以类似于 LINPACK/LAPACK 的方式将这个聚类框架实现为一组库函数,其他研究人员可以使用它更有效地构建自己的聚类功能; (d) 通过聚类分析提供对若干生物学问题的见解; (e) 使用我们的聚类框架作为训练场,培训学生/博士后如何构建生物数据分析工具。我们框架的基础是数据集及其与聚类关系的最小生成树(MST)表示。我们的初步研究表明:(i)MST 与聚类概念之间存在天然联系,有助于将多维数据聚类问题简化为树划分问题; (ii) 可以最优且高效地解决具有在(最小生成)树上定义的一般目标函数的聚类问题; (iii) MST 提供了一个自然的框架来解决更一般的聚类问题,即从噪声背景中提取数据聚类。其他初步研究还表明,MST 具有与聚类相关的丰富特性,进一步的研究可能会带来更加有效的聚类和分析生物数据的方法。我们的研究将分为五个任务来组织和开展。 o 研究 MST 与聚类的基本属性:我们将研究 MST 与聚类之间的基本关系。关于它们之间关系的新见解和发现将为开发更有效的聚类方法奠定基础。 o 基于 MST 的聚类算法和统计分析方法的研究和开发:我们将针对几个聚类相关问题研究和开发一大类基于 MST 的算法。此外,我们将研究和开发有效的统计分析工具,以评估聚类结果的统计显着性和稳健性。 o 为四个选定的应用问题开发改进的分析能力:我们将把我们的聚类框架应用于四个生物数据分析问题:(1)基因表达数据分析,(2)调节结合位点识别,(3)二杂交数据分析,和(4)系统发育树聚类分析。 o 将我们基于 MST 的聚类框架实现为库函数:我们将实现我们的基于 MST 的聚类框架聚类相关的算法作为 API(应用程序编程接口),其他研究人员可以在自己的数据分析软件中轻松使用。此外,我们将把我们的聚类工具实现为社区服务的 Web 服务器。 o 培训和教育:由于 MST 提供了如此丰富的与聚类相关的有吸引力的属性,我们将使用我们基于 MST 的聚类框架作为培训平台,教学生/博士后如何开发生物数据分析工具。我们提出的研究和开发直接解决 ITR 计划在以下领域的研究挑战: o 提供新的计算、模拟和数据分析方法和工具来建模物理、生物、社会、行为和数学现象,并改进我们理解、建模和控制复杂系统行为的能力。
英文摘要
Project SummaryThe Human Genome Project has opened the flood-gate of biological data, which has resulted in the generation of enormous amount of sequence, structure, expression, and interaction data at rates that far exceed our current capability of analyzing and interpreting them. New ideas and approaches are urgently needed to establish greatly improved capabilities for biological data analysis. Data clustering is fundamental to mining a large quantity of biological data. The goals of this project are (a) to develop a highly effective and general framework for biological data clustering, which is applicable to a large class of biological data analysis problems; (b) to demonstrate the effectiveness of this framework as a general-purpose clustering tool, through application to four challenging biological data analysis problems; (c) to implement this clustering framework as a set of library functions, in a similar fashion to LINPACK/LAPACK, with which other researchers can build their own clustering capabilities more efficiently; (d) to provide insight on several biological problems through clustering analysis; and (e) to train students/postdocs how to build biological data analysis tools, using our clustering framework as a training ground. The foundation of our framework is a minimum spanning tree (MST) representation of a data set and its relationships with clustering. Our preliminary studies have revealed that (i) there is a natural connection between MSTs and the concept of clustering, which can help to reduce a multi-dimensional data clustering problem to a tree-partitioning problem; (ii) clustering problems with general objective functions, defined on (minimum spanning) trees, can be solved optimally and efficiently; and (iii) MSTs provide a natural framework for solving a more general class of clustering problems, i.e., extracting data clusters from a noisy background. Additional preliminary studies have also revealed that MSTs have such rich properties related to clustering that further investigation could lead to significantly more effective ways of clustering and analyzing biological data. Our research will be organized and carried out in five tasks.o Investigation of fundamental properties of MSTs versus clustering: We will investigate fundamental relationships between MSTs and clustering. New insights and discoveries about their relationships will be used to lay the foundation for development of more effective ways of clustering.o Investigation and development of MST-based clustering algorithms and statistical analysis methods: We will investigate and develop a large class of MST-based algorithms for several clustering related problems. In addition, we will investigate and develop effective statistical analysis tools for assessing statistical significance and robustness of clustering results.o Development of improved analysis capabilities for four selected application problems: We will apply our clustering framework to four biological data analysis problems: (1) gene expression data analysis, (2) regulatory binding site identification, (3) two-hybrid data analysis, and (4) phylogenetic tree clustering analysis.o Implementation of our MST-based clustering framework as library functions: We will implement our MST-based clustering-related algorithms as APIs (Application Programming Interface), which can be used easily by other researchers in their own data analysis software. In addition, we will implement our clustering tools as a Web server for community service.o Training and education: As MST provides such a rich set of attractive properties relevant to clustering, we will use our MST-based clustering framework as a training platform to teach students/postdocs how to develop biological data analysis tools.Our proposed study and development directly address the research challenges of the ITR program in the following areas:o providing new computational, simulation and data-analysis methods and tools to model physical, biological,social, behavioral and mathematical phenomena, ando improving our ability to understand, model and control the behavior of complex systems.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Building A Teacher-AI Collaborative System for Personalized Instruction and Assessment of Comprehension Skills
-
批准号:2302730
-
项目类别:Standard Grant
-
资助金额:$85.0万
-
财政年份:2023
-
负责人:Ying Xu
-
依托单位:
UNS: Organophosphates and Phthalates in Sleep Microenvironments: Emission, Transport, and Infants' Exposure
-
批准号:1512610
-
项目类别:Continuing Grant
-
资助金额:$35.08万
-
财政年份:2015
-
负责人:Ying Xu
-
依托单位:
CAREER: Emission and Transport of PBDEs in Indoor Environments
-
批准号:1150713
-
项目类别:Standard Grant
-
资助金额:$40.91万
-
财政年份:2012
-
负责人:Ying Xu
-
依托单位:
Collaborative Research: Phthalate Plasticizers: Temperature Dependence of Material/Air Equilibria and Consequences for Emissions, Exposure and Risk
-
批准号:1066642
-
项目类别:Continuing Grant
-
资助金额:$15.0万
-
财政年份:2011
-
负责人:Ying Xu
-
依托单位:
MRI: Acquisition of a Computer Cluster for Bioinformatics Research at UGA
-
批准号:0821263
-
项目类别:Standard Grant
-
资助金额:$79.68万
-
财政年份:2008
-
负责人:Ying Xu
-
依托单位:
Computational Prediction of Biological Networks in Microbes and Applications to Cyanobacteria
-
批准号:0542119
-
项目类别:Continuing Grant
-
资助金额:$187.58万
-
财政年份:2006
-
负责人:Ying Xu
-
依托单位:
CompBio: A New Paradigm of Protein Threading: simultaneous backbone threading and side-chain packing prediction.
-
批准号:0621700
-
项目类别:Standard Grant
-
资助金额:$27.5万
-
财政年份:2006
-
负责人:Ying Xu
-
依托单位:
A Computational Capability for Fast and Reliable Characterization of Protein Complexes
-
批准号:0354771
-
项目类别:Continuing Grant
-
资助金额:$0.0万
-
财政年份:2003
-
负责人:Ying Xu
-
依托单位:
ITR Collaborative Research: Combinatorial Algorithms for Biological Data Clustering
-
批准号:0407204
-
项目类别:Continuing Grant
-
资助金额:$129.5万
-
财政年份:2003
-
负责人:Ying Xu
-
依托单位:
A Computational Capability for Fast and Reliable Characterization of Protein Complexes
-
批准号:0213840
-
项目类别:Continuing Grant
-
资助金额:$88.86万
-
财政年份:2002
-
负责人:Ying Xu
-
依托单位:
海外基金