Hidden Hamiltonian Cycle Recovery via Linear Programming

Hidden Hamiltonian Cycle Recovery via Linear Programming
复制标题

DOI:
10.1287/opre.2019.1886
复制
发表时间:
2018-04
期刊:
Oper. Res.
影响因子:
--
通讯作者:
V. Bagaria;Jian Ding;David Tse;Yihong Wu;Jiaming Xu
V. Bagaria;Jian Ding;David Tse;Yihong Wu;Jiaming Xu
中科院分区:
其他
文献类型:
--
作者:
V. Bagaria;Jian Ding;David Tse;Yihong Wu;Jiaming Xu

文献摘要

相似文献

我们介绍了隐藏的哈密顿循环恢复的问题,有一个未知的哈密顿循环在一个$n$-顶点完全图,需要从噪声边缘测量推断。测量是独立的,并根据循环中的边缘的$\卡尔普_n$和$\calQ_n$分布。该公式是由基因组组装中的问题激发的,其中目标是使用重叠群之间的长距离连接测量根据它们在基因组上的位置对一组重叠群(基因组重叠群)进行排序。在这个模型中计算最大似然估计减少到一个旅行商问题(TSP)。尽管TSP的NP-困难,我们表明,一个简单的线性规划(LP)松弛,即分数$2$-因子(F2 F)LP,恢复隐藏的哈密顿循环具有高概率为$n \to \infty$提供$\alpha_n - \log n \to \infty$,其中$\alpha_n \triangleq-2\log \int \sqrt{d P_n d Q_n}$是阶数为$\frac{1}{2}$的R\'enyi散度。这个条件是信息理论上最优的,在温和的分布假设下,$\alpha_n \geq(1+o(1))\log n$是任何算法成功所必需的,而不管计算成本如何。从通常的证明技术的基础上的双重见证建设,分析依赖于组合特征(特别是,半完整性)的极值点的F2 F多面体。这些极值点被表示为双色多重图,进一步分解为更简单的“花型”结构,用于大偏差分析和计数参数。在真实的数据上的算法评估表明了对现有方法的改进。
We introduce the problem of hidden Hamiltonian cycle recovery, where there is an unknown Hamiltonian cycle in an $n$-vertex complete graph that needs to be inferred from noisy edge measurements. The measurements are independent and distributed according to $\calP_n$ for edges in the cycle and $\calQ_n$ otherwise. This formulation is motivated by a problem in genome assembly, where the goal is to order a set of contigs (genome subsequences) according to their positions on the genome using long-range linking measurements between the contigs. Computing the maximum likelihood estimate in this model reduces to a Traveling Salesman Problem (TSP). Despite the NP-hardness of TSP, we show that a simple linear programming (LP) relaxation, namely the fractional $2$-factor (F2F) LP, recovers the hidden Hamiltonian cycle with high probability as $n \to \infty$ provided that $\alpha_n - \log n \to \infty$, where $\alpha_n \triangleq -2 \log \int \sqrt{d P_n d Q_n}$ is the R\'enyi divergence of order $\frac{1}{2}$. This condition is information-theoretically optimal in the sense that, under mild distributional assumptions, $\alpha_n \geq (1+o(1)) \log n$ is necessary for any algorithm to succeed regardless of the computational cost. Departing from the usual proof techniques based on dual witness construction, the analysis relies on the combinatorial characterization (in particular, the half-integrality) of the extreme points of the F2F polytope. Represented as bicolored multi-graphs, these extreme points are further decomposed into simpler "blossom-type" structures for the large deviation analysis and counting arguments. Evaluation of the algorithm on real data shows improvements over existing approaches.