CIF: Small: Collaborative Research: Error Correction with Natural Redundancy
CIF: Small: Collaborative Research: Error Correction with Natural Redundancy
批准号:
1717884
负责人:
Jehoshua Bruck
金额:
$16.67万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-08-01 至 2021-07-31
中文摘要
第1部分:项目的非技术描述该项目研究通过使用数据的内部结构从数据中去除错误的基本问题。这表明,当前数据存储系统中存储的大量数据具有非常丰富的结构;因此,充分利用它们进行纠错,可以显著提高数据存储系统的可靠性。该项目研究了该技术的几个基本方面,包括如何发现和表征各种类型数据的高度复杂结构,如何利用它们有效地纠正数据中的错误以提高数据存储系统的可靠性,如何将该技术与现有的基于向数据添加外部冗余的纠错技术相结合,以及如何在实际数据存储系统中实现该技术。该项目解决了现代社会的一个关键问题:如何确保数据能够大规模、长时间地可靠存储。这项新技术有可能大大提高信息基础设施的可靠性,这些基础设施经常为科学和工业计算访问大量数据。该项目本质上是跨学科的:它结合了多个科学领域,包括信息理论、机器学习、大数据分析和算法设计,旨在教育学生并为下一代存储系统的劳动力发展做出贡献。该项目结合了严谨的理论分析和重要的实际应用,以促进学术界和工业界之间的合作,并在共同努力下创造新的科学进步。第2部分:项目技术描述本项目研究如何利用大数据固有的冗余进行纠错。大数据的例子包括语言、图像、数据库等。内置冗余与纠错码(ECC)相结合,有效纠错。目标是将存储系统中的数据可靠性提升到一个新的水平。为了实现这一目标,将开发新的技术来发现压缩和未压缩数据中各种类型的固有冗余。将探索结合固有冗余解码器和ECC解码器的新方法,以实现有效的纠错。将研究容量和计算复杂性的基本限制,以便使用固有冗余进行纠错。这个项目结合了纠错和机器学习,本质上是跨学科的。它将从多个方面扩展现有的纠错知识。首先,它使用自然语言处理和深度学习技术来发现大数据中适合纠错的新型冗余,这些冗余超出了当前联合源信道编码的知识范围。这包括针对已经被各种压缩算法压缩的数据的冗余发现技术。其次,它探索了ecc的解码算法,不仅具有规则的ecc强加冗余,而且具有不规则的固有冗余。它扩展了现有的纠错方案,从容量和计算复杂性两方面对纠错的固有冗余进行了基本限制。第三,将理论研究与实际系统相结合,为下一代大数据存储和传输系统奠定基础。现代社会越来越依赖于数字数据。随着每天产生的数据爆炸式增长,必须在纠错方面取得进展,以赶上数据爆炸式增长的速度。该项目旨在将数据可靠性显著提高到一个新的水平,这一方向的改进对现代社会的日常工作和生活非常有益。这个项目是编码理论和机器学习之间的跨学科,可以促进信息理论和计算机科学社区之间的合作。该项目将严谨的理论分析与重要的实际应用相结合,促进学术界和工业界的合作,共同创造新的科学进步。拟议的研究将通过为研究生和本科生开发新课程,并让代表性不足的国内和国际学生参与高级研究,与工程教育相结合。研究结果将在国家/国际会议和期刊上积极宣传。
英文摘要
Part 1: Nontechnical description of the projectThis project studies the fundamental problem of removing errors from data by using internal structures of data. It shows that the vast amount of data stored in current data-storage systems possess very rich structures; therefore, by fully exploiting them for error correction, the reliability of data-storage systems can be improved significantly. The project studies several fundamental aspects of this technology, including how to discover and characterize the highly complex structures of various types of data, how to use them to correct errors in data efficiently to improve the reliability of data-storage systems, how to combine the technology with existing error-correction techniques that are based on adding external redundancy to data, and how to implement the technology in practical data-storage systems. This project addresses a critical issue of the modern society: how to ensure that data can be stored reliably at large scale and over a long time. The new technology has the potential to substantially improve the dependability of information infrastructure, which accesses vast amounts of data frequently for scientific and industrial computing. The project is interdisciplinary in nature: it combines multiple scientific fields including information theory, machine learning, big data analysis and algorithm design, and aims to educate students and contribute to workforce development for next-generation storage systems. The project conjugates rigorous theoretical analysis and significant practical applications, to foster collaboration between academia and industry, and create new scientific advances with combined efforts.Part 2: Technical description of the projectThis project studies how to use the inherent redundancy in big data for error correction. Examples of big data include languages, images, databases, and others. The inherent redundancy is integrated with error-correcting codes (ECC) for effective error correction. The objective is to elevate data reliability in storage systems to the next level. To achieve this goal, new techniques will be developed to discover various types of inherent redundancy in both compressed and uncompressed data. New approaches will be explored to combine inherent-redundancy decoders and ECC decoders for effective error correction. Fundamental limits of both capacity and computational complexity will be studied for error correction using inherent redundancy.This project combines error correction with machine learning and is interdisciplinary in nature. It will expand the current knowledge on error correction in multiple ways. First, it uses techniques in natural language processing and deep learning to discover new types of redundancy in big data that are suitable for error correction, and which extend beyond current knowledge in joint source-channel coding. This includes redundancy discovery techniques for data already compressed by various compression algorithms. Second, it explores decoding algorithms for ECCs with not only regular ECC-imposed redundancy, but also irregular inherent redundancy. It extends existing error correction schemes to cast the fundamental limits of inherent redundancy for error correction, in terms of both capacity and computational complexity. Third, by integrating a theoretical study with practical systems, a foundation can be laid for next-generation systems that store and transmit big data.Modern society relies increasingly heavily on digital data. With the explosive amount of data generated each day, it is essential to make advances in error correction that can catch the speed of data explosion. This project aims at improving data reliability significantly to the next level, and improvements in this direction can be highly beneficial to the daily work and life of the modern society. This project, being interdisciplinary between coding theory and machine learning, can foster collaboration between the information theory and computer science communities. The project combines rigorous theoretical analysis with significant practical applications, to foster collaboration between academia and industry, and create new scientific advances with combined efforts. The proposed research will be integrated with engineering education by developing new courses for graduate and undergraduate students, and involving under-represented, domestic and international students in advanced research. The results will be actively publicized in national/international conferences and journals.
期刊论文(13)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1109/isit44484.2020.9174311
发表时间:
2020-01
期刊:
2020 IEEE International Symposium on Information Theory (ISIT)
影响因子:
--
作者:
[Netanel Raviv;Siddhartha Jain;Jehoshua Bruck]
通讯作者:
Netanel Raviv;Siddhartha Jain;Jehoshua Bruck
Two Deletion Correcting Codes from Indicator Vectors
指示向量的两个删除校正代码
DOI:
10.1109/isit.2018.8437868
发表时间:
2020
期刊:
IEEE transactions on information theory
影响因子:
2.5
作者:
[Sima, J., Raviv, N. and]
通讯作者:
Raviv, N. and
DOI:
10.1109/isit.2019.8849547
发表时间:
2019-01
期刊:
2019 IEEE International Symposium on Information Theory (ISIT)
影响因子:
--
作者:
[Netanel Raviv;Qian Yu;Jehoshua Bruck;A. Avestimehr]
通讯作者:
Netanel Raviv;Qian Yu;Jehoshua Bruck;A. Avestimehr
DOI:
10.1109/isit.2019.8849783
发表时间:
2019
期刊:
International Symposium on Information Theory and its Applications
影响因子:
--
作者:
[Sima, J., Bruck, J.]
通讯作者:
Bruck, J.
Evolution of $k$ -Mer Frequencies and Entropy in Duplication and Substitution Mutation Systems
复制和替换突变系统中 $k$ -Mer 频率和熵的演化
DOI:
10.1109/tit.2019.2946846
发表时间:
2020
期刊:
IEEE Transactions on Information Theory
影响因子:
2.5
作者:
[Lou, Hao, Schwartz, Moshe, Bruck, Jehoshua, Farnoud, Farzad]
通讯作者:
Farnoud, Farzad
共 9 条
CIF: NSF-BSF: Small: Collaborative Research: Characterization and Mitigation of Noise in a Live DNA Storage Channel
-
批准号:1816965
-
项目类别:Standard Grant
-
资助金额:$18.73万
-
财政年份:2018
-
负责人:Jehoshua Bruck
-
依托单位:
CIF: Small: Collaborative Research: Coding for Green Storage Technologies in Nonvolatile Memories
-
批准号:1218005
-
项目类别:Standard Grant
-
资助金额:$25.0万
-
财政年份:2012
-
负责人:Jehoshua Bruck
-
依托单位:
Collaborative Research: BRAM: Balanced RAnk Modulation for data storage in next generation flash memories
-
批准号:0801795
-
项目类别:Standard Grant
-
资助金额:$15.0万
-
财政年份:2008
-
负责人:Jehoshua Bruck
-
依托单位:
NR: Collaborative Research: Scheduling for Efficient and Reliable Data Broadcast
-
批准号:0322475
-
项目类别:Continuing Grant
-
资助金额:$15.79万
-
财政年份:2003
-
负责人:Jehoshua Bruck
-
依托单位:
Collaborative Research: Efficient Data Distribution Schemes for Secure and Reliable Networked Storage Systems
-
批准号:0209042
-
项目类别:Standard Grant
-
资助金额:$12.0万
-
财政年份:2002
-
负责人:Jehoshua Bruck
-
依托单位:
NYI: Efficient Fault-Tolerant Parallel and Distributed Computing
-
批准号:9457811
-
项目类别:Continuing Grant
-
资助金额:$31.25万
-
财政年份:1994
-
负责人:Jehoshua Bruck
-
依托单位:
国内基金
海外基金
登录
查看更多内容
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:
-
依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2022
-
负责人:张祥忠
-
依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
-
批准号:32000033
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:林平
-
依托单位:
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
-
批准号:31972324
-
项目类别:面上项目
-
资助金额:58.0万元
-
批准年份:2019
-
负责人:高学文
-
依托单位:
变异链球菌small RNAs连接LuxS密度感应与生物膜形成的机制研究
-
批准号:81900988
-
项目类别:青年科学基金项目
-
资助金额:21.0万元
-
批准年份:2019
-
负责人:毛梦莹
-
依托单位:
肠道细菌关键small RNAs在克罗恩病发生发展中的功能和作用机制
-
批准号:31870821
-
项目类别:面上项目
-
资助金额:56.0万元
-
批准年份:2018
-
负责人:陈江宁
-
依托单位:
基于small RNA 测序技术解析鸽分泌鸽乳的分子机制
-
批准号:31802058
-
项目类别:青年科学基金项目
-
资助金额:26.0万元
-
批准年份:2018
-
负责人:麻慧
-
依托单位:
Small RNA介导的DNA甲基化调控的水稻草矮病毒致病机制
-
批准号:31772128
-
项目类别:面上项目
-
资助金额:60.0万元
-
批准年份:2017
-
负责人:吴建国
-
依托单位:
基于small RNA-seq的针灸治疗桥本甲状腺炎的免疫调控机制研究
-
批准号:81704176
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2017
-
负责人:赵继梦
-
依托单位:
水稻OsSGS3与OsHEN1调控small RNAs合成及其对抗病性的调节
-
批准号:91640114
-
项目类别:重大研究计划
-
资助金额:85.0万元
-
批准年份:2016
-
负责人:何祖华
-
依托单位: