课题基金 / 基金详情

CIF: Small: Collaborative Research: Inference of Information Measures on Large Alphabets: Fundamental Limits, Fast Algorithms, and Applications

CIF: Small: Collaborative Research: Inference of Information Measures on Large Alphabets: Fundamental Limits, Fast Algorithms, and Applications
CIF:小型:协作研究:大字母表上信息测量的推断:基本限制、快速算法和应用
批准号:
1528159
负责人:
Tsachy Weissman
金额:
$25.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-09-01 至 2018-08-31

项目摘要

项目成果

Tsachy Weissman的其他基金

相似基金

相关文献

中文摘要
翻译
信息论中的一项关键任务是描述压缩、通信和更一般的涉及信息存储、传输和处理的操作问题的基本性能极限。这样的刻画通常是关于信息度量的,其中最基本的是香农熵和互信息。除了在传统信息论领域中的突出操作角色外,信息度量还在许多统计建模和机器学习任务中找到了大量的应用。各种现代数据分析应用程序处理自然被视为大范围内概率分布的样本的数据集。由于典型的大字母表大小和资源限制,从业者在从语料库语言学到神经科学的各种应用中都面临着采样不足的困难。该项目的主要目标之一是在一套新的数学工具的基础上发展一种一般理论,这将促进对大字母表上的信息计量的最佳估计的构建和分析。该项目的另一个主要方面是将新的理论方法纳入机器学习算法,从而显著影响当前的现实世界学习实践。该项目的成功完成将导致能够实现的技术和实用方案--从神经反应数据的分析到学习图形模型的应用--被证明比现有的更接近于达到基本性能极限。该项目的发现将丰富现有的大数据分析课程。将开设一门专门研究高维统计推断的新课程,深入探讨大字母表数据的估计问题。将在斯坦福大学和UIUC组织和举办关于该项目主题和调查结果的讲习班。通过基于最佳多项式逼近的计算效率的程序,以及可证明的基本最优性保证,将开发一种全面的逼近理论方法来估计大字母表上分布的泛函。植根于高维统计文献,我们的关键观察是,虽然估计分布本身需要样本大小与字母表大小线性缩放,但有可能以次线性样本复杂性精确估计分布的泛函,如熵或互信息。这需要超越传统智慧,开发比最大可能性(插件)更复杂的方法。估计。该项目的另一个主要方面是将新的理论方法转化为高度可扩展和高效的机器学习算法,从而显著影响当前的真实世界学习实践,并显著提高几个最流行的机器学习应用程序的性能,例如学习依赖于互信息估计的图形模型。
英文摘要
A key task in information theory is to characterize fundamental performance limits in compression, communication, and more general operational problems involving the storage, transmission and processing of information. Such characterizations are usually in terms of information measures, among the most fundamental of which are the Shannon entropy and the mutual information. In addition to their prominent operational roles in the traditional realms of information theory, information measures have found numerous applications in many statistical modeling and machine learning tasks. Various modern data-analytic applications deal with data sets naturally viewed as samples from a probability distribution over a large domain. Due to the typically large alphabet size and resource constraints, the practitioner contends with the difficulty of undersampling in applications ranging from corpus linguistics to neuroscience. One of the main goals of this project is the development of a general theory based on a new set of mathematical tools that will facilitate the construction and analysis of optimal estimation of information measures on large alphabets. The other major facet of this project is the incorporation of the new theoretical methodologies into machine learning algorithms, thereby significantly impacting current real-world learning practices. Successful completion of this project will result in enabling technologies and practical schemes - in applications ranging from analysis of neural response data to learning graphical models - that are provably much closer to attaining the fundamental performance limits than existing ones. The findings of this project will enrich existing big data-analytic curricula. A new course dedicated to high-dimensional statistical inference that addresses estimation for large-alphabet data in depth will be created and offered. Workshops on the themes and findings of this project will be organized and held at Stanford and UIUC. A comprehensive approximation-theoretic approach to estimating functionals of distributions on large alphabets will be developed via computationally efficient procedures based on best polynomial approximation, with provable essential optimality guarantees. Rooted in the high-dimensional statistics literature, our key observation is that while estimating the distribution itself requires the sample size to scale linearly with the alphabet size, it is possible to accurately estimate functionals of the distribution, such as entropy or mutual information, with sub-linear sample complexity. This requires going beyond the conventional wisdom by developing more sophisticated approaches than maximal likelihood (?plug-in?) estimation. The other major facet of this project is translating the new theoretical methodologies into highly scalable and efficient machine learning algorithms, thereby significantly impacting current real-world learning practices and significantly boosting the performance in several of the most prevalent machine learning applications, such as learning graphical models, that rely on mutual information estimation.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: CIF: Medium: An Information-Theoretic Foundation for Adaptive Bidding in First-Price Auctions
  • 批准号:
    2106467
  • 项目类别:
    Standard Grant
  • 资助金额:
    $45.0万
  • 财政年份:
    2021
  • 负责人:
    Tsachy Weissman
  • 依托单位:
CIF:Small:Collaborative Research: Compressed databases for similarity queries: fundamental limits and algorithms
  • 批准号:
    1321174
  • 项目类别:
    Standard Grant
  • 资助金额:
    $25.0万
  • 财政年份:
    2013
  • 负责人:
    Tsachy Weissman
  • 依托单位:
EAGER: Action in Information Processing
  • 批准号:
    1049413
  • 项目类别:
    Standard Grant
  • 资助金额:
    $30.0万
  • 财政年份:
    2010
  • 负责人:
    Tsachy Weissman
  • 依托单位:
Collaborative Research: The Role of Feedback in Two-Way Communication Networks
  • 批准号:
    0729119
  • 项目类别:
    Standard Grant
  • 资助金额:
    $30.0万
  • 财政年份:
    2007
  • 负责人:
    Tsachy Weissman
  • 依托单位:
国内基金
海外基金
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
  • 依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2022
  • 负责人:
    张祥忠
  • 依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
  • 批准号:
    31972324
  • 项目类别:
    面上项目
  • 资助金额:
    58.0万元
  • 批准年份:
    2019
  • 负责人:
    高学文
  • 依托单位: