CRII: CIF: New Paradigms in Generalization and Information-Theoretic Analysis of Deep Neural Networks
CRII: CIF: New Paradigms in Generalization and Information-Theoretic Analysis of Deep Neural Networks
批准号:
1947801
负责人:
Ziv Goldfeld
金额:
$17.5万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-04-01 至 2023-03-31
中文摘要
在过去的十年中,深度学习(DL)已经成为各种机器学习任务的首选方法。DL应用领域不断扩展,现在包括自动驾驶汽车,机器人辅助手术,医学成像等。社会对这些技术的广泛接受取决于人类理解和信任它们的能力。不幸的是,深度学习系统的特殊实际效果并没有与一个全面的理论相结合,来解释它们是如何运作的,以及为什么它们在现实世界的数据上如此成功。这种状况阻碍了上述应用程序更广泛地部署AI。为了缓解这一僵局,该项目试图打开深度神经网络(DNN)的引擎盖,使深度学习成为可能,并阐明信息在这些系统中是如何处理的。这样做将使人工智能机制的决策对最终用户和其他利益相关者更加透明,从而有助于他们的理解。通过严格的性能保证,该项目还旨在描述深度学习系统不会失败的情况。这些进步将为高性能人工智能系统在我们日常生活中的整合奠定基础,释放其宝贵的潜在影响。 该项目通过一种新的信息理论方法来解决DL理论中的关键挑战。主要目标是阐明DNN逐步构建表示的过程-从浅层的粗糙和过度冗余表示到深层的高度聚类和可解释表示-并让设计师对该过程进行更多控制。为此,将努力实现三个协同增效的目标。首先是通过量化DNN中的信息流来开发内部表示的新复杂性度量。至关重要的是,这些措施是为了有效地计算层的维度典型的最先进的网络,用于计算机视觉,语音和文本处理。第二个推力的重点是通过新的实例相关的泛化范围的网络的泛化能力的复杂性措施。这里的目标是为给定的DNN提供性能保证,以有效地计算品质因数。最后,进一步利用所开发的机器来构建用于修剪冗余神经元/层、可视化DNN的操作以及提高DNN可解释性的工具。总而言之,这项研究致力于将DNN设计目前不确定的试错过程推向确定性工程实践领域。该奖项反映了NSF的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Over the past decade, deep learning (DL) has become the method of choice for various machine learning tasks. The realm of DL applications constantly expands, now including autonomous vehicles, robotic-assisted surgery, medical imaging, and many others. A wide societal acceptance of such technologies relies on the ability of humans to understand and trust them. Unfortunately, the exceptional practical effectiveness of DL systems is not coupled with a comprehensive theory to explain how they operate and why they are so successful on real-world data. This state of affairs obstructs a wider deployment of AI for the applications described above. To alleviate this impasse, this project seeks to open the hood of Deep Neural Networks (DNNs) that enable DL and elucidate how information is processed in these systems. Doing so would make the decisions of AI mechanisms more transparent to end users and other stakeholders, thus contributing to their understanding. Via rigorous performance guarantees, this project also aims to characterize the circumstances under which deep learning system are warranted not to fail. These advances will set the stage for the integration of high-performance AI systems in our daily lives, unlocking their invaluable potential impact. The project tackles key challenges in DL theory via a novel information-theoretic approach. The main objective is to shed light on the process by which DNNs progressively build representations --- from crude and over-redundant representations in shallow layers, to highly-clustered and interpretable ones in deeper layers --- and to give the designer more control over that process. To that end, three synergistic thrusts are pursued. First is developing novel complexity measures of internal representations by quantifying the flow of information through the DNN. Crucially, these measures are designed for efficient computation over layer dimensionalities typical to state-of-the-art networks for computer vision, speech, and text processing. The second thrust focuses on relating the developed complexity measures to the generalization capability of the network via new instance-dependent generalization bounds. The goal here is to provide performance guarantees for a given DNN in terms of efficiently computable figures of merit. Lastly, the developed machinery is further leveraged to construct tools for pruning redundant neurons/layers, visualizing the DNN's operation, and progressing DNN interpretability. Altogether, this research strives to progress the current uncertain trial-and-error process of DNN design towards the domain of deterministic engineering practice.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(4)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Capacity of Continuous Channels with Memory via Directed Information Neural Estimator
通过定向信息神经估计器存储连续通道的容量
DOI:
10.1109/isit44484.2020.9174109
发表时间:
2020
期刊:
IEEE International Symposium on Information Theory
影响因子:
--
作者:
[Aharoni, Ziv, Tsur, Dor, Goldfeld, Ziv, Permuter, Haim H.]
通讯作者:
Permuter, Haim H.
DOI:
10.48550/arxiv.2206.08526
发表时间:
2022-06
期刊:
ArXiv
影响因子:
--
作者:
[Ziv Goldfeld;K. Greenewald;Theshani Nuradha;Galen Reeves]
通讯作者:
Ziv Goldfeld;K. Greenewald;Theshani Nuradha;Galen Reeves
DOI:
--
发表时间:
2022
期刊:
Journal of machine learning research
影响因子:
6
作者:
[Sreekumar, Sreejith, Goldfeld, Ziv]
通讯作者:
Goldfeld, Ziv
DOI:
--
发表时间:
2022
期刊:
IEEE International Symposium on Information Theory
影响因子:
--
作者:
[D. Tsur, Z. Aharoni]
通讯作者:
D. Tsur, Z. Aharoni
NSF-BSF: Collaborative Research: CIF: Small: Neural Estimation of Statistical Divergences: Theoretical Foundations and Applications to Communication Systems
-
批准号:2308446
-
项目类别:Standard Grant
-
资助金额:$20.0万
-
财政年份:2023
-
负责人:Ziv Goldfeld
-
依托单位:
CAREER: Smooth statistical distances for a scalable learning theory
-
批准号:2046018
-
项目类别:Continuing Grant
-
资助金额:$64.18万
-
财政年份:2021
-
负责人:Ziv Goldfeld
-
依托单位:
国内基金
海外基金
Wolbachia的cif因子与天麻蚜蝇dsx基因协同调控生殖不育的机制研究
-
批准号:JCZRQN202501187
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2025
-
负责人:
-
依托单位:
SHR和CIF协同调控植物根系凯氏带形成的机制
-
批准号:31900169
-
项目类别:青年科学基金项目
-
资助金额:23.0万元
-
批准年份:2019
-
负责人:李朋雪
-
依托单位: