Low-dimensional representation of genomic sequences

Low-dimensional representation of genomic sequences
复制标题

DOI:
10.1007/s00285-019-01348-1
复制
发表时间:
2019-03
影响因子:
1.9
通讯作者:
Richard C. Tillquist;M. Lladser
Richard C. Tillquist;M. Lladser
中科院分区:
数学4区
文献类型:
--
作者:
Richard C. Tillquist;M. Lladser

文献摘要

相似文献

许多数据分析和数据挖掘技术需要将数据嵌入到欧几里得空间中。当面对符号数据集时,特别是由高通量测序测定产生的生物序列数据时,诸如二进制和k-mer计数向量的常规嵌入方法可能维度太高或粒度太粗而不能有效地从数据中学习。其他表示技术,如多维缩放(MDS)和Node 2 Vec可能不适合大型数据集,因为它们需要在面对新的未分类数据时从头开始重新计算完整的嵌入。为了克服这些问题,我们修改了图论的概念“度量维数”的“简化”。就像三边测量可以用来表示欧几里得平面上的点,通过它们到三个非共线点的距离来表示点一样,三角测量允许我们通过它到节点子集的距离来表示图中的任何节点。不幸的是,确定最小子集和最低维嵌入的问题对于一般图来说是NP完全的。然而,通过专门研究特别适合表示生物序列的汉明图,我们可以很容易地生成低维嵌入,将任意长度的序列映射到真实的空间。作为概念验证,我们使用MDS,Node 2 Vec和基于多侧化的嵌入对以内含子-外显子边界为中心的DNA 20聚体进行分类。尽管这些不同的技术执行重复测序,但MDS和Node 2 Vec潜在地遭受随着序列长度增加的可扩展性问题,而重复测序提供了映射长基因组序列的有效手段。
Numerous data analysis and data mining techniques require that data be embedded in a Euclidean space. When faced with symbolic datasets, particularly biological sequence data produced by high-throughput sequencing assays, conventional embedding approaches like binary andk-mer count vectors may be too high dimensional or coarse-grained to learn from the data effectively. Other representation techniques such as Multidimensional Scaling (MDS) and Node2Vec may be inadequate for large datasets as they require recomputing the full embedding from scratch when faced with new, unclassified data. To overcome these issues we amend the graph-theoretic notion of “metric dimension” to that of “multilateration.” Much like trilateration can be used to represent points in the Euclidean plane by their distances to three non-colinear points, multilateration allows us to represent any node in a graph by its distances to a subset of nodes. Unfortunately, the problem of determining a minimal subset and hence the lowest dimensional embedding is NP-complete for general graphs. However, by specializing to Hamming graphs, which are particularly well suited to representing biological sequences, we can readily generate low-dimensional embeddings to map sequences of arbitrary length to a real space. As proof-of-concept, we use MDS, Node2Vec, and multilateration-based embeddings to classify DNA 20-mers centered at intron–exon boundaries. Although these different techniques perform comparably, MDS and Node2Vec potentially suffer from scalability issues with increasing sequence length whereas multilateration provides an efficient means of mapping long genomic sequences.