Collaborative Research: A General Framework for High Throughput Biological Learning: Theory Development and Applications
Collaborative Research: A General Framework for High Throughput Biological Learning: Theory Development and Applications
批准号:
0714669
负责人:
Shaw-Hwa Lo
金额:
$27.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2007
资助国家:
美国
项目状态:
已结题
起止时间:
2007-09-15 至 2011-08-31
中文摘要
该应用程序提供了一个全面的研究计划,用于研究处理生物(医学)和其他科学研究产生的复杂大规模数据集的一般框架和各种新方法。这项提议明确了两个目标:理论发展和在生物学和医学中的应用。前者专注于研究一个通用的、核心的、无模型的框架,以有效地解决高维数据产生的主要问题。在后者中,研究人员试图应用从理论部分发展起来的方法来解决生物和医学中出现的机器学习类型的问题。特别是,该团队打算研究与治疗响应的生物学和医学预测、疾病(如癌症)的临床诊断、蛋白质-蛋白质相互作用的发现以及与疾病病因和基序识别相关的生物网络构建等问题。为了实现这两个目标,研究人员将在一般情况下研究理论和实用性质,并评估一系列新的统计/计算程序/软件,然后将通过广泛的真实和模拟数据进行测试,其中一些数据来自当前正在进行的研究。成功地处理低维数据的方法对于高维数据不再有效。分析这些数据的最大困难之一是识别信息性变量/特征及其关联的集群,并破译这些变量和集群之间相互作用的特征。为了满足当前和未来从高维数据中全面、系统地挖掘隐藏知识的需求,科学领域必须开发新的方法。目前的项目是对这一需求的直接回应。基于在提取低维信息方面已经获得的理论证据(作为初步结果),该团队计划应用并开发各种有效的程序来解决生物学和医学领域中的实际重要问题。研究人员将研究一种适用于多个领域的新的筛选过程,以演示如何在充分利用影响变量之间的联合信息的同时识别出高质量的低维分类器。为了进一步解释生物验证/确认,该团队将研究如何基于低维分类器构建生物网络,以及如何识别它们之间的重要关联模式。将在方法学开发小组和生物验证小组之间建立反馈机制,定期讨论统计/计算结果并对其进行生物验证。预计这里开发的关键思想和方法将在生物/医学以外的学科中得到大量应用。拟议的研究可能会促进实质性的知识,并极大地有利于分子生物学/统计学/计算生物学/疾病预测/药物发现方面当前和未来的努力。该项目还将为本科生提供宝贵的研究经验和培训。
英文摘要
This application presents a comprehensive research plan for the investigation of a general framework and various new methods to handle complex large-scale data sets generated from biological (medical) as well as other scientific studies. Two goals are articulated in this proposal: theory development and application in biology and medicine. The former is focused on the study of a general yet core, model-free framework to effectively address major issues arising from high dimensional data. In the latter, the investigators seek to apply methods developed from the theory part to resolve machine learning type problems that arise in biology and medicine. In particular, this team intends to study the problems related to biological and medical prediction in response to treatments, clinical diagnosis of diseases (such as cancers), discovery of protein-protein interactions and biological network constructions related to disease etiology and motif identification. To achieve these two goals, the investigators will study theoretical and practical properties under a general setting and evaluate a series of novel statistical/computation procedures/software which will then be tested by a broad range of real and simulated data, some from current on-going studies.The emergence of high dimensional data in most scientific fields poses new challenges for statisticians. Methods successful in dealing with low dimensional data are no longer effective for high dimensional data. One of the greatest difficulties in analyzing these data is to identify the informative variables/features and their associated clusters, and decipher the characteristics of the interaction between these variables and clusters. To meet current and future needs for digging hidden knowledge out of high dimensional data comprehensively and systematically, the scientific fields must develop new methods. The current project is a direct response to this need. Based on theoretical evidence (as preliminary results) already obtained in extracting low dimensional information, this team plans to apply and to develop various effective procedures to address practically important problems in the domains of biology and medicine. The investigators will study a novel screening process applicable across fields to demonstrate how high quality classifiers of low dimensionality can be identified while joint information among the influential variables are fully utilized. For further interpretation for biological validation/confirmation this team will study how to construct biological networks based on low dimensional classifiers and how to identify significant association patterns among them. A feedback mechanism will be established between the methodology development and biological validation teams, where statistical/computational results will be regularly discussed and biologically validated. It is anticipated that the key ideas and methods developed here will find numerous applications in disciplines other than biology/medicine. The proposed research is likely to advance substantial knowledge and significantly benefit current and future efforts in molecular biology/statistics/computational biology/disease prediction/drug discovery. The project would also provide valuable research experiences and training to undergraduates.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
BIGDATA: F: Statistical Foundation of Predictivity: A Novel Architecture for Big Data Learning
-
批准号:1741191
-
项目类别:Standard Grant
-
资助金额:$90.0万
-
财政年份:2018
-
负责人:Shaw-Hwa Lo
-
依托单位:
A Novel Statistical Framework for Big Data Prediction
-
批准号:1513408
-
项目类别:Standard Grant
-
资助金额:$30.0万
-
财政年份:2015
-
负责人:Shaw-Hwa Lo
-
依托单位:
Statistical Analysis of Linkage/Association on Family-Based Studies in Human Genetics
-
批准号:0071930
-
项目类别:Continuing Grant
-
资助金额:$26.05万
-
财政年份:2000
-
负责人:Shaw-Hwa Lo
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Research on Quantum Field Theory without a Lagrangian Description
-
批准号:24ZR1403900
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:SATOSHI NAWATA
-
依托单位:
Cell Research
-
批准号:31224802
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2012
-
负责人:程磊
-
依托单位:
Cell Research
-
批准号:31024804
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2010
-
负责人:程磊
-
依托单位:
Cell Research (细胞研究)
-
批准号:30824808
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2008
-
负责人:张爱兰
-
依托单位:
Research on the Rapid Growth Mechanism of KDP Crystal
-
批准号:10774081
-
项目类别:面上项目
-
资助金额:45.0万元
-
批准年份:2007
-
负责人:滕冰
-
依托单位: