Efficient Encoding/Decoding of Irreducible Words for Codes Correcting Tandem Duplications

Efficient Encoding/Decoding of Irreducible Words for Codes Correcting Tandem Duplications
复制标题

不可约字的高效编码/解码,用于纠正串联重复的代码

DOI:
10.1109/isit.2018.8437789
复制
发表时间:
2018
期刊:
2018 IEEE International Symposium on Information Theory (ISIT)
影响因子:
--
通讯作者:
T. T. Nguyen
T. T. Nguyen
中科院分区:
--
文献类型:
--
作者:
Yeow Meng Chee;Johan Chrisnata;Han Mao Kiah;T. T. Nguyen

文献摘要

被引文献

相似文献

串联复制是将 DNA 片段的副本插入到原始位置附近的过程。 Jain 等人受到在生物体中存储数据的应用程序的启发。 (2017) 提出了纠正串联重复的代码研究。所有代码构造都基于不可约词。我们研究不可约词的有效编码/解码方法。首先,我们描述一个 $(\ell,\ m)$ 有限状态编码器,并表明当 $m=\Theta(1/\epsilon)$ 和 $\ell=\Theta(1/\epsilon)$ 时,编码器的速率与最优值相差 $\epsilon$。接下来,我们提供不可约词的排序/取消排序算法,并修改算法以减少有限状态编码器的空间需求。
Tandem duplication is the process of inserting a copy of a segment of DNA adjacent to the original position. Motivated by applications that store data in living organisms, Jain et al. (2017) proposed the study of codes that correct tandem duplications. All code constructions are based on irreducible words. We study efficient encoding/decoding methods for irreducible words. First, we describe an $(\ell,\ m)$ -finite state encoder and show that when $m=\Theta(1/\epsilon)$ and $\ell=\Theta(1/\epsilon)$, the encoder has rate that is $\epsilon$ away from the optimal. Next, we provide ranking/unranking algorithms for irreducible words and modify the algorithms to reduce the space requirements for the finite state encoder.