课题基金 / 基金详情

BIGDATA: F: Towards Automating Data Analysis: Interpretable, Interactive, and Scalable Learning via Discrete Probability

BIGDATA: F: Towards Automating Data Analysis: Interpretable, Interactive, and Scalable Learning via Discrete Probability
BIGDATA:F:迈向自动化数据分析:通过离散概率进行可解释、交互式和可扩展的学习
批准号:
1741341
负责人:
Suvrit Sra
金额:
$102.44万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-10-01 至 2022-09-30

项目摘要

项目成果

Suvrit Sra的其他基金

相似基金

相关文献

中文摘要
翻译
随着机器学习(ML)渗透到科学和技术的各个领域,不同数据领域、推理问题、资源限制和可靠性的需求引发了一些新的概念和算法挑战。目前限制充分使用机器学习的缺点的例子包括:数据和算法的使用不够理想;费力的手工调整和模型搜索;结果的验证和推广的困难;与人类的互动有限;领域知识编码;以及缺乏可解释性等等。在这些问题上的进展有可能影响机器学习在广泛领域的成功采用和使用。在上述动机下,该项目的目标是创建一套新的模型和算法来分析复杂的数据集,特别关注以下三个对下一代机器学习至关重要的因素:(1)可解释性;(2)互动性;(3)自动学习。这一提议背后的主要技术概念是离散概率中的负依赖概念。该项目为以这一概念为基础的一套新工具奠定了理论基础。除了实际影响外,项目中研究的方法还会激发新的理论问题,并有助于提高人们对基础数学的兴趣。拟议工作的实际影响可能会在多个方面造福社会。通过合作,PI将评估在医疗保健(寻求最终影响患者护理和福祉)、系统生物学(帮助癌症和糖尿病等方面的研究)和材料科学(帮助更有效地发现更安全、功能材料)方面开发的方法。该项目还将直接产生教育影响:培训研究生,为各级数据科学课程提供材料,并通过一般性讲座以及在会议和讲习班上的重点讲座,包括针对数据科学领域妇女的讲习班和活动,向社区进行宣传。从技术上讲,PIS将开发:(1)用于交互数据分析的新工具、模型和算法,特别是用于实验设计、信息收集、可解释的机器学习、假设检验、性能验证和体系结构学习;(2)理论分析,如收敛和复杂性(统计和计算);以及(3)所有关键算法和框架的开源实现。
英文摘要
As machine learning (ML) permeates all areas of science and technology, demands in diverse data domains, inference questions, resource limitations and reliability fuel several new conceptual and algorithmic challenges. Examples of current shortcomings that limit the full use of machine learning include suboptimal use of data and algorithms; painstaking hand-tuning and model search; validation of results and difficulties in generalization; limited interactivity with humans; encoding of domain knowledge; and lack of interpretability, among others. Progress on these questions has the potential to impact the successful adoption and use of machine learning in a broad range of fields. With the above motivation, the goal of this project is to create a novel suite of models and algorithms for analyzing complex datasets, with a particular focus on the following three factors crucial for next-generation machine learning: (1) interpretability; (2) interactivity; and (3) automated learning. The overarching technical concept underlying this proposal is the concept of negative dependence in discrete probability. This project lays theoretical foundations for a new set of tools grounded in this concept. Besides practical impacts, the methods to be studied in the project motivate new theoretical questions, and will help increase interest in the underlying mathematics.The practical impact of the proposed work has the potential to benefit society on multiple fronts. Via collaborations, the PIs will evaluate the developed methods in healthcare (seeking to ultimately impact patient care and well-being), systems biology (to help with research on cancer and diabetes, among others), and materials science (to help discover safer, functional materials more efficiently). The project will also directly have educational impact: training of graduate students, providing material for data science courses at all levels, and outreach to the community via general talks as well as focused lectures at conferences and workshops, including workshops and events targeted at women in Data Science. Technically, the PIs will develop: (1) New tools, models, and algorithms for interactive data analysis, especially for experimental design, information collection, interpretable machine learning, hypothesis testing, performance validation, and architecture learning; (2) Theoretical analysis, such as convergence and complexity (statistical and computational); and (3) Open-source implementations of all key algorithms and frameworks.
期刊论文(21)
专著(0)
科研奖励(0)
会议论文
Provably Efficient Algorithms for Multi-Objective Competitive RL
可证明有效的多目标竞争强化学习算法
DOI: --
发表时间: 2021
期刊: Proceedings of the 38th International Conference on Machine Learning
影响因子: --
作者: [Yu, Tiancheng, Tian, Yi, Zhang, Jingzhao, Sra, Suvrit]
通讯作者: Sra, Suvrit
Discrete Sampling using Semigradient-based Product Mixtures
使用基于半梯度的产品混合物进行离散采样
DOI: --
发表时间: 2018
期刊: UAI 2018 Proceedings
影响因子: --
作者: [Gotovos, Alkis, Hassani, Hamed, Krause, Andreas, Jegelka, Stefanie]
通讯作者: Jegelka, Stefanie
DOI: --
发表时间: 2021-06
期刊: ArXiv
影响因子: --
作者: [Ching-Yao Chuang;Youssef Mroueh;K. Greenewald;A. Torralba;S. Jegelka]
通讯作者: Ching-Yao Chuang;Youssef Mroueh;K. Greenewald;A. Torralba;S. Jegelka
DOI: 10.1609/aaai.v36i8.20796
发表时间: 2021-12
期刊: ArXiv
影响因子: --
作者: [Anshul B. Shah;S. Sra;Ramalingam Chellappa;A. Cherian]
通讯作者: Anshul B. Shah;S. Sra;Ramalingam Chellappa;A. Cherian
共 16 条
    CAREER: Modern nonconvex optimization for machine learning: foundations of geometric and scalable techniques
    TRIPODS+X:RES:Collaborative Research: Learning with Expert-In-The-Loop for Multimodal Weakly Labeled Data and an Application to Massive Scale Medical Imaging
    海外基金