课题基金 / 基金详情

CRII: III: Robust Machine Learning Methods for Messy Data

CRII: III: Robust Machine Learning Methods for Messy Data
CRII:III:针对杂乱数据的鲁棒机器学习方法
批准号:
1657155
负责人:
James Zou
金额:
$17.5万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-05-01 至 2019-08-31

项目摘要

项目成果

James Zou的其他基金

相似基金

相关文献

中文摘要
翻译
杂乱的数据在现代科学中无处不在。数据来自不同的来源;有许多潜在的混杂因素;而且人们常常不清楚该问哪些相关的问题,该使用哪些模型。这一现实与机器学习和统计学通常的建模假设形成鲜明对比,在这些假设中,数据被假设来自明确指定的模型,需要测试的假设也被明确列出。标准理论和混乱数据的实际实践之间的明显差距是导致科学界再现性危机的主要原因,也阻碍了研究人员充分利用数据的洞察力。该项目将开发严格的数学基础和强大的机器学习算法,以解决混乱数据的核心挑战。PI将探索新的技术来量化和减少探索性数据分析产生的不同类型的选择偏差。PI还将研究在模型错误或指定不足时执行统计推断的算法。该项目将应用这些新方法来解决人类基因组学中具有挑战性的问题。PI最近初始化了一个基于信息使用的框架,以量化数据探索引起的过度拟合和偏差的程度。该项目将大大扩展这一框架。特别是PI将应用这种信息使用方法来量化和减少自适应实验产生的数据偏差,例如在线A/B测试和更一般的多臂盗。与过拟合相关的是统计模型的错误和指定不足的问题。PI最近开发了累积方法,用于在观测受到未知和任意干扰时学习概率模型。一个有希望的研究方向是将这种方法扩展到允许非线性干扰的更一般的设置中,并为广泛的数据科学社区开发软件工具。基因组学体现了许多混乱数据的挑战——基因组数据通常需要大量的探索性分析,并面临建模的不确定性。这使得基因组学成为应用这里开发的新的凌乱数据算法的高影响领域。生物医学数据库是由许多研究人员交互分析的,因此特别容易出现探索偏差和过拟合。该项目将探索在斯坦福大学创建的生物医学数据中心试行信息使用框架,以量化和减少探索偏差。作为该项目的一部分,PI还开发了课程、研讨会和教程,将机器学习、统计学、信息论和生物医学数据科学领域的研究人员和实践者聚集在一起,以解决混乱数据带来的无处不在的挑战。
英文摘要
Messy data is ubiquitous in modern science. Data come from heterogeneous sources; there are many latent confounding factors; and it is often unclear what are the relevant questions to ask and models to use. This reality is in sharp contrast with the usual modeling assumptions of machine learning and statistics, where data are assumed to come from well-specified models and the hypotheses to test are clearly laid out. The glaring gap between standard theory and the actual practice of messy data is a major contributor to the reproducibility crises across science and prevents researchers from harnessing the full insights from data. This project will develop rigorous mathematical foundations and robust machine learning algorithms to address the core challenges of messy data. The PI will explore novel techniques to quantify and reduce different types of selection biases that arise from exploratory data analysis. The PI will also investigate algorithms to perform statistical inference when the model is mis- or under-specified. The project will apply these new methods to tackle challenging problems in human population genomics. The PI recently initialized a framework based on information usage to quantify the magnitude of over-fitting and bias arising from data exploration. This project will significantly expand this framework. In particular the PI will apply this information usage approach to quantify and reduce bias in data generated from adaptive experimentation, such as online A/B testing and more general multi-arm bandits. Related to over-fitting is the problem of mis- and under-specified statistical models. The PI has recently developed method-of-cumulant approaches to learn probabilistic models when the observations are perturbed by unknown and arbitrary interference. A promising direction of research is to extend this approach to more general settings that allow for nonlinear interference and to develop software tools for the broad data science community. Genomics exemplify many of the challenges of messy data-genomic data typically requires substantial exploratory analysis and faces modeling uncertainty. This makes genomics a high impact domain to apply the new messy data algorithms developed here. Bio-medical databases are interactively analyzed by many researchers and thus are particularly prone to exploration bias and overfitting. The PI will explore piloting the information usage framework on the bio-medical data hubs being created at Stanford in order to quantify and reduce exploration bias. As a part of the project, PI is also developing courses, workshops and tutorials to bring together researchers and practitioners across machine learning, statistics, information theory and bio-medical data science to address the ubiquitous challenge of messy data.
期刊论文(5)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1038/s41588-018-0295-5
发表时间: 2019-01-01
期刊: NATURE GENETICS
影响因子: 30.8
作者: [Zou, James, Huss, Mikael, Telenti, Amalio]
通讯作者: Telenti, Amalio
DOI: 10.1038/s41467-018-04608-8
发表时间: 2018-05-30
期刊: Nature communications
影响因子: 16.6
作者: [Abid A, Zhang MJ, Bagaria VK, Zou J]
通讯作者: Zou J
NeuralFDR: learning decision threshold from hypothesis features.
NeuralFDR:从假设特征中学习决策阈值。
DOI: --
发表时间: 2017
期刊: Proceedings of Neural Information Processing Conference
影响因子: --
作者: [Xia, Fei, Zhang, Martin, Zou, James, Tse, David]
通讯作者: Tse, David
DOI: 10.1073/pnas.1720347115
发表时间: 2018-04-17
期刊: PROCEEDINGS OF THE NATIONAL ACADEMY OF SCIENCES OF THE UNITED STATES OF AMERICA
影响因子: 11.1
作者: [Garg, Nikhil, Schiebinger, Londa, Zou, James]
通讯作者: Zou, James
CAREER: Enabling data valuation and deletion in human-centered machine learning
  • 批准号:
    1942926
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $49.99万
  • 财政年份:
    2020
  • 负责人:
    James Zou
  • 依托单位:
AF: MEDIUM: Collaborative Research: Foundations of Adaptive Data Analysis
  • 批准号:
    1763191
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $27.6万
  • 财政年份:
    2018
  • 负责人:
    James Zou
  • 依托单位:
国内基金
海外基金
基于人工智能与多组学的III期结核性脓胸CT“低密度线”形成机制及手术时机预测模型研究
基于MOF–CRISPR微流控平台的雄黄As(III)/As(V)价态识别与炮制耦合机制研究
  • 批准号:
    JCZRLH202600780
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
  • 依托单位:
白术内酯III靶向IRF4-CD36轴通过调控脂质代谢重编程提升结直肠癌奥沙利铂敏感性的机制研究
  • 批准号:
    2026JJ82690
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
    张卓
  • 依托单位:
基于废水零排放的FeS-As(III)置换法从污酸中清洁脱砷处理技术研究
  • 批准号:
    2026JJ30130
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
    张二军
  • 依托单位: