课题基金 / 基金详情

CAREER: Flexible Parsimonious Models for Complex Data

CAREER: Flexible Parsimonious Models for Complex Data
职业:复杂数据的灵活简约模型
批准号:
1653017
负责人:
Jacob Bien
金额:
$40.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-07-01 至 2017-08-31

项目摘要

项目成果

Jacob Bien的其他基金

相似基金

相关文献

中文摘要
翻译
学术界、工业界和政府的研究人员正在以远远超出之前想象的规模和复杂程度生成数据。复杂的数据需要足够灵活的统计模型来适应有意义的潜在信号,使科学家能够发现意想不到的模式。然而,随着社会越来越多地依赖统计算法来做出影响日常生活的决策,非专家能够解释一种方法的输出变得越来越重要。这需要简约:比起复杂的解释,更简单的解释更受青睐。例如,互联网以文本的形式产生了前所未有的海量数据(如文章、博客、网页、消费者评论和许多其他社交媒体产品)。这些文本数据代表了洞察世界的潜在宝库--人们在想什么,这种情况随着时间的推移正在发生什么变化,这种变化如何因地点而异,等等。研究人员开发了新的统计方法,以克服主要的技术挑战,从这些数据中收集有用的信息。同样的方法论也可以应用于微生物群的研究,微生物群是生活在人类肠道等环境中的巨大微生物群落。需要更好的统计方法来确定肠道中对人类健康和疾病起关键作用的微生物类型。该项目解决的另一个问题涉及对随时间收集的数据(如风速数据和野生动物监测)进行建模。开发的方法允许更准确的预测,这在许多领域至关重要,包括卫生和医药以及开发成本较低的能源系统。该项目的最后一个主要领域致力于使统计研究过程更有效率,使其软件质量更高,更容易在统计研究人员群体中共享。最后,所有这三个研究目标都与教育成果紧密结合在一起,包括研究生的监督和教学,对非统计学家和非科学家的推广,以及发布本科生可访问的迷你论文,描述研究人员的新研究成果。该项目侧重于设计新的统计方法,以平衡两个重要且往往相互矛盾的需求:灵活性和简洁性。(1)当特征高度稀疏时,建立预测回归和分类模型是困难的。虽然许多方法关注的是高维的挑战,但相对较少的方法考虑了很少是非零的特征所构成的障碍。研究人员开发了一种新的特征选择框架,当特征高度稀疏时,该框架在现有方法失败的情况下成功。从理论和计算两个角度对此进行了研究。(2)高维协方差估计和时间序列建模是统计学中两个丰富但基本不同的领域,研究者将这两个领域结合起来,发展了局部平稳时间序列建模的新方法。从稳定到地方稳定增加的灵活性必须谨慎地与节俭相平衡。(3)将在调查员的新平台上免费分发一系列针对特定领域的软件模块,以简化进行模拟研究的过程。每个模块都将实现在统计研究的特定领域中使用的一些最常见的模型、方法和指标。其目标是通过创建一种易于适应的标准化格式,促进在统计研究界共享高质量、可重复使用的模拟代码。
英文摘要
Researchers throughout academia, industry, and government are generating data at scales and levels of complexity far beyond what could previously have been imagined. Complex data demand statistical models that are sufficiently flexible to adapt to meaningful, underlying signals, allowing scientists to discover unexpected patterns. Yet as society relies more heavily on statistical algorithms to make decisions impacting everyday life, it becomes increasingly important for a method's output to be interpretable by non-experts. This demands parsimony: that simpler explanations be favored over more complicated ones. For example, the Internet has led to unprecedented quantities of data in the form of text (such as articles, blogs, webpages, consumer reviews, and many other social media products). Such text data represent a potential treasure trove of insights into the world -- what people are thinking, how this is changing over time, how this varies by location, etc. The investigator develops new statistical methods for overcoming major technical challenges to gleaning useful information from this data. This same methodology can be applied to the study of the microbiome, the vast community of microbes living in an environment such as the human gut. Better statistical methods are needed to identify types of microbes in the gut that play a crucial role in human health and disease. Another problem that is tackled in this project involves modeling data collected over time (such as wind-speed data and wildlife monitoring). The methods that are developed allow for more accurate forecasting, which is crucial in many areas including health and medicine and the development of lower cost energy systems. The last major area in this project is devoted to making the process of statistical research more efficient and its software of higher quality and easier to share across the community of statistical researchers. Finally, all three research objectives are closely integrated with educational outcomes, including the supervision and teaching of graduate students, outreach to non-statisticians and non-scientists, and the release of undergraduate-accessible mini-papers describing the investigator's new research findings.This project focuses on the design of new statistical methods that balance two important and often opposing needs: flexibility and parsimony. (1) Building predictive regression and classification models is difficult when the features are highly sparse. While many methods focus on the challenge of high dimensionality, relatively few have considered the obstacle posed by features that are rarely nonzero. The investigator develops a new framework for feature selection when the features are highly sparse that succeeds where preexisting methods fail. This is studied both from theoretical and computational standpoints. (2) High-dimensional covariance estimation and time series modeling are two rich, but largely distinct, areas in statistics, which the investigator combines to develop new methods for modeling locally stationary time series. The added flexibility in going from stationarity to local stationarity must be carefully balanced with parsimony. (3) A series of area-specific software modules will be distributed freely online building on the investigator's new platform for streamlining the process of performing simulation studies. Each module will implement some of the most common models, methods, and metrics used in a given area of statistics research. The goal is to facilitate the sharing of high-quality, reproducible simulation code in the statistics research community by creating an easily-adaptable standardized format.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CAREER: Flexible Parsimonious Models for Complex Data
  • 批准号:
    1748166
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $40.0万
  • 财政年份:
    2017
  • 负责人:
    Jacob Bien
  • 依托单位:
High-Dimensional Covariance Estimation via Convex Optimization
  • 批准号:
    1405746
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $12.0万
  • 财政年份:
    2014
  • 负责人:
    Jacob Bien
  • 依托单位:
国内基金
海外基金
A study on prototype flexible multifunctional graphene foam-based sensing grid (柔性多功能石墨烯泡沫传感网格原型研究)
  • 批准号:
    --
  • 项目类别:
    --
  • 资助金额:
    20万元
  • 批准年份:
    2020
  • 负责人:
    SAGAR RIZWAN UR REHMAN
  • 依托单位: