CAREER: Statistical Methodology in Multi-view Learning with Large Data
CAREER: Statistical Methodology in Multi-view Learning with Large Data
批准号:
1255045
负责人:
Mark Culp
金额:
$40.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2013
资助国家:
美国
项目状态:
已结题
起止时间:
2013-06-01 至 2018-05-31
中文摘要
在复杂的科学研究分类和回归问题中,通常使用几个不同的数据集来描述响应。这些数据集中的每一个都提供了响应的唯一视图,但通常这些视图都不能完美地描述响应。研究者开发了计算效率高的统计方法,在统计分析、变量和视图选择以及结果解释的框架内对多视图数据进行建模。该方法既基于正则化方法,也基于加法模型中的稀疏惩罚和基于算法的密集迭代方法。此外,研究者为残差分析提供了坚实的基础,建立了模型的一致性,并解决了合纵性的实际问题。研究者致力于提高科学研究界和行业对使用这些建模技术分析涉及多个来源数据的复杂研究问题的优势的认识,并通过这些数据集的实践经验培训未来的统计学家和相关专业人员。为了支持这项工作,研究者开发并维护了一个功能强大、用户友好的统计软件包来实施这一方法。在现代世界中,我们从许多不同来源收集数据的能力已经大大扩展,部分原因是过去几十年计算机的创新。没有跟上的是同时分析来自许多不同来源的数据的能力。因此,学术界和工业界的科学研究人员没有充分利用可以通过适当组合多个不同来源的数据以提供可解释结果的方式找到的信息。这是一个具有挑战性的问题,涉及到统计学、计算机科学、数学和数据库管理方面的进展。研究者从统计分析的角度来解决这个问题。这项研究的许多应用涉及新的统计方法,以帮助癌症研究、药理学、遗传学、蛋白质组学、文本数据处理和国土安全。在当今的科学研究中,对大型、复杂、多视图数据集的分析需求是巨大的,而这一需求目前尚未得到满足。这项工作的结果正在改变许多领域的研究人员分析和解释数据的方式。
英文摘要
In complex scientific research classification and regression problems, it is common for several different data sets to be used to describe the response. Each of these data sets provides a unique view of the response, but typically none of these views describes the response perfectly. The investigator develops computationally efficient statistical methodology to model multi-view data within a framework for statistical analysis, variable and view selection, and the interpretation of results. The methodology is based on both regularization approaches involving sparse penalties in additive models and algorithmic intensive iterative-based approaches. In addition, the investigator provides a solid foundation for analysis of residuals in this context, establishes the consistency of the model, and addresses the practical issue of concurvity. The investigator is committed to raising awareness in scientific research communities and industries of the advantages of using these modeling techniques in the analysis of complex research problems involving data from multiple sources and to training future statisticians and related professionals through hands-on experiences with these data sets. To support this effort, the investigator develops and maintains a powerful, user-friendly statistical software package to implement this methodology.In the modern world our ability to collect data from many different sources has expanded dramatically due in part to computer innovations over the past few decades. What has not kept pace is the ability to analyze data from many different sources simultaneously. As a result, scientific researchers in academia and industry are not fully harnessing the information that can be found by appropriately combining multiple, diverse sources of data in a way that can provide interpretable results. This is a challenging problem involving advances in statistics, computer science, mathematics, and database management. The investigator addresses this problem from a statistical analysis viewpoint. Many applications of this research involve new statistical methods to help with cancer research, pharmacology, genetics, proteomics, text data processing, and homeland security. The need for analyses of large, complex, multi-view data sets is substantial in scientific research today, and this need is currently unmet. The results of this work are transforming how researchers in many fields analyze and interpret data.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金