Approximate parameterized matching

Approximate parameterized matching
复制标题

近似参数化匹配

DOI:
--
复制
发表时间:
2004
期刊:
TALG
影响因子:
--
通讯作者:
Dina Sokol
Dina Sokol
中科院分区:
--
文献类型:
--
作者:
Carmit Hazay;Moshe Lewenstein;Dina Sokol

文献摘要

被引文献

相似文献

两个相等长度的字符串<i>s</i>和<i>s</i>′,在字母表s和s ′上,<i>参数匹配,</i>如果存在双射π:<sub><i>s</i></sub>→<sub><i>s</i></sub> ′使得π(<i>s</i>)=<i>s</i> ′,其中π(<i>s</i>)是<i>s</i>的每个字符通过π的重命名。<sub><i></i></sub><sub><i></i></sub><i>参数化匹配</i>是在文本t中找到模式串p的所有参数化匹配的问题<i></i><i></i>,而<i>近似参数化匹配</i>是在每个位置找到最大化从p映射到适当字符串的字符数的双射π的问题<i></i>。|<i>p</i>|- 长度子串<i>t</i>。 参数化匹配作为软件维护系统中的软件重复检测模型被引入,并且在图像处理和计算生物学中也有应用。例如,近似参数化匹配在存在误差的情况下对具有可变颜色图的图像搜索进行建模。 我们考虑给定误差阈值k的问题<i></i>,目标是找到t中<i></i>存在双射π的所有位置,该双射π将p映射<i></i>到适当的|<i>p</i>|- 长度的子串<i></i>,最多有<i>k</i>个不匹配的映射元素。我们的主要结果是这个问题的一个算法,时间复杂度为<i>O</i>(<i>nk</i><sup>1.5</sup> + <i></i>nlogm<i></i>),其中<i>m</i> =|<i>p</i>|和<i>n</i> =|<i>不</i>|.我们还表明,当|<i>p</i>| = |<i>不</i>|= <i>m</i>,这个问题等价于图上的最大匹配问题,产生<i>O</i>(<i>m</i> + <i>k</i><sup>1.5</sup>)的解。
Two equal length strings <i>s</i> and <i>s</i>′, over alphabets Σ<sub><i>s</i></sub> and Σ<sub><i>s</i></sub>′, <i>parameterize match</i> if there exists a bijection π : Σ<sub><i>s</i></sub> → Σ<sub><i>s</i></sub>′ such that π (<i>s</i>) = <i>s</i>′, where π (<i>s</i>) is the renaming of each character of <i>s</i> via π. <i>Parameterized matching</i> is the problem of finding all parameterized matches of a pattern string <i>p</i> in a text <i>t</i>, and <i>approximate parameterized matching</i> is the problem of finding at each location a bijection π that maximizes the number of characters that are mapped from <i>p</i> to the appropriate |<i>p</i>|-length substring of <i>t</i>. Parameterized matching was introduced as a model for software duplication detection in software maintenance systems and also has applications in image processing and computational biology. For example, approximate parameterized matching models image searching with variable color maps in the presence of errors. We consider the problem for which an error threshold, <i>k</i>, is given, and the goal is to find all locations in <i>t</i> for which there exists a bijection π which maps <i>p</i> into the appropriate |<i>p</i>|-length substring of <i>t</i> with at most <i>k</i> mismatched mapped elements. Our main result is an algorithm for this problem with <i>O</i>(<i>nk</i><sup>1.5</sup> + <i>mk</i> log <i>m</i>) time complexity, where <i>m</i> = |<i>p</i>| and <i>n</i>=|<i>t</i>|. We also show that when |<i>p</i>| = |<i>t</i>| = <i>m</i>, the problem is equivalent to the maximum matching problem on graphs, yielding a <i>O</i>(<i>m</i> + <i>k</i><sup>1.5</sup>) solution.