课题基金 / 基金详情

CAREER: Flexible Network Estimation from High-Dimensional Data

CAREER: Flexible Network Estimation from High-Dimensional Data
职业:根据高维数据进行灵活的网络估计
批准号:
1252624
负责人:
Daniela Witten
金额:
$40.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2013
资助国家:
美国
项目状态:
已结题
起止时间:
2013-07-01 至 2020-06-30

项目摘要

项目成果

Daniela Witten的其他基金

相似基金

相关文献

中文摘要
翻译
这项研究涉及新的统计方法和理论的发展,图形建模的基础上,高维数据,其中的功能数量超过了观察。在某些应用中,例如基于基因表达数据的转录调控网络的估计,现有技术由于两个原因而不足:这些技术所基于的假设在面对如此高的维度时不足以进行准确的网络恢复,此外,所做的假设对于数据可能是不现实的。为了解决这两个问题,研究者提出研究(a)通过对真实条件依赖网络的拓扑结构进行更有效和结构化的假设,通过凸罚和其他技术,更有效地学习一个或多个高斯图模型的一组技术;以及(B)用于估计条件依赖关系而无需通常的高斯假设的更灵活的框架。近年来,新技术和快速的计算机已经导致在分子生物学、市场营销、金融、社会学、语言学和计算机视觉等不同领域中产生和获得大量数据。不幸的是,分析这种类型的“大数据”带来了严峻的统计挑战,经典的统计工具集无法应用。因此,开发有效的统计机器学习技术来理解非常大规模的数据集对于许多科学和工业领域的进步至关重要,以便弥合正在收集的数据与关于数据的科学和工业问题之间的差距。例如,能够根据基因组数据估计基因网络对于理解生物过程以及在治疗癌症和其他疾病方面取得进展具有重要意义。该提案包括:(1)开发基于高维数据集的改进网络估计技术;(2)通过出版物、研讨会和公开发布软件向统计和生物医学界传播所产生的技术;(3)培训博士生掌握大数据的统计机器学习技术;以及(4)通过短期课程、会议演示和其他活动,增加高中生、本科生和代表性不足群体的成员对统计机器学习和大数据挑战的接触。
英文摘要
This research involves the development of new statistical methods and theory for graphical modeling on the basis of high-dimensional data, in which the number of features exceeds the number of observations. In certain applications, such as the estimation of transcriptional regulatory networks on the basis of gene expression data, existing techniques are inadequate for two reasons: the assumptions that underlie these techniques are insufficient for accurate network recovery in the face of such high dimensionality, and furthermore the assumptions that are made may be unrealistic for the data. To address these two problems, the investigator proposes to study (a) a set of techniques for more effectively learning one or more Gaussian graphical models by making more effective and structured assumptions about the topology of the true conditional dependence networks, via convex penalties and other techniques; and (b) more flexible frameworks for estimating conditional dependence relationships without the usual Gaussianity assumptions.In recent years, new technologies and fast computers have resulted in the generation and availability of vast amounts of data in fields as diverse as molecular biology, marketing, finance, sociology, linguistics, and computer vision. Unfortunately, analyzing this type of "big data" poses severe statistical challenges, and the classical statistical toolset cannot be applied. Therefore, developing effective statistical machine learning techniques for making sense of very large-scale data sets is crucial for progress in many areas of science as well as industry, in order to bridge the gap between the data that is being collected and the scientific and industrial questions that are being asked about the data. As an example, being able to estimate gene networks on the basis of genomic data has important implications for understanding biological processes, and for making progress towards the treatment of cancer and other disease. This proposal involves (1) developing techniques for improved network estimation on the basis of high-dimensional data sets; (2) disseminating the resulting techniques to the statistical and biomedical communities via publications, seminars, and the public release of software; (3) training PhD students in statistical machine learning techniques for big data; and (4) increasing the exposure of high school students, undergraduates, and members of underrepresented groups to statistical machine learning and big data challenges via short courses, conference presentations, and other activities.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CDS&E-MSS: Data thinning: methods, theory, and applications
  • 批准号:
    2322920
  • 项目类别:
    Standard Grant
  • 资助金额:
    $25.0万
  • 财政年份:
    2023
  • 负责人:
    Daniela Witten
  • 依托单位:
国内基金
海外基金
A study on prototype flexible multifunctional graphene foam-based sensing grid (柔性多功能石墨烯泡沫传感网格原型研究)
  • 批准号:
    --
  • 项目类别:
    --
  • 资助金额:
    20万元
  • 批准年份:
    2020
  • 负责人:
    SAGAR RIZWAN UR REHMAN
  • 依托单位: