Robust k-mer frequency estimation using gapped k-mers.

Robust k-mer frequency estimation using gapped k-mers.
复制标题

DOI:
10.1007/s00285-013-0705-3
复制
发表时间:
2014-08
影响因子:
1.9
通讯作者:
Beer, Michael A.
Beer, Michael A.
中科院分区:
数学4区
文献类型:
--
作者:
Ghandi, Mahmoud;Mohammad-Noori, Morteza;Beer, Michael A.

文献摘要

参考文献

被引文献

相似文献

固定长度k的寡聚体,通常称为k-mer,通常用作描述不同生物学功能的DNA序列特征的基本元件,或用作构建更复杂的序列特征描述符(如位置权重矩阵)的中间元件。k-mer作为一般序列特征是非常有用的,因为它们构成完整和无偏的特征集,并且不需要基于生物学机制的不完全知识的参数化。然而,使用k-mer作为序列特征的基本限制是,随着k的增加,可以描述DNA序列元件中更大的空间相关性,但是观察到任何特定k-mer的频率变得非常小,并且迅速接近二进制计数的稀疏矩阵。因此,一旦k变大,使用k-mer的任何统计学习方法将容易受到k-mer频率的噪声估计的影响。由于所有的分子DNA相互作用具有有限的空间范围,缺口的k-mer通常携带相关的生物信号。在这里,我们使用gapped k-mer计数来更稳健地估计unapped k-mer频率,通过推导出给定观察到的一组gapped k-mer频率的k-mer频率的最小范数估计的方程。我们证明,这种方法提供了一个更准确的估计的k-mer频率在真实的生物序列中使用的样本CTCF结合位点在人类基因组中。
Oligomers of fixed length, k, commonly known as k-mers, are often used as fundamental elements in the description of DNA sequence features of diverse biological function, or as intermediate elements in the constuction of more complex descriptors of sequence features such as position weight matrices. k-mers are very useful as general sequence features because they constitute a complete and unbiased feature set, and do not require parameterization based on incomplete knowledge of biological mechanisms. However, a fundamental limitation in the use of k-mers as sequence features is that as k is increased, larger spatial correlations in DNA sequence elements can be described, but the frequency of observing any specific k-mer becomes very small, and rapidly approaches a sparse matrix of binary counts. Thus any statistical learning approach using k-mers will be susceptible to noisy estimation of k-mer frequencies once k becomes large. Because all molecular DNA interactions have limited spatial extent, gapped k-mers often carry the relevant biological signal. Here we use gapped k-mer counts to more robustly estimate the ungapped k-mer frequencies, by deriving an equation for the minimum norm estimate of k-mer frequencies given an observed set of gapped k-mer frequencies. We demonstrate that this approach provides a more accurate estimate of the k-mer frequencies in real biological sequences using a sample of CTCF binding sites in the human genome.
DOI: 10.1093/bioinformatics/16.1.16
发表时间: 2000-01-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Stormo, GD
通讯作者: Stormo, GD
DOI: 10.1093/bioinformatics/btg425
发表时间: 2004-02-12
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
van Helden, J
通讯作者: van Helden, J
DOI: 10.1101/gr.112656.110
发表时间: 2011-03-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Boyle, Alan P.;Song, Lingyun;Furey, Terrence S.
通讯作者: Furey, Terrence S.
支持向量机和计算生物学的内核。
DOI: 10.1371/journal.pcbi.1000173
发表时间: 2008-10
影响因子: 4.3
作者:
Ben-Hur, Asa;Ong, Cheng Soon;Sonnenburg, Soeren;Schoelkopf, Bernhard;Raetsch, Gunnar
通讯作者: Raetsch, Gunnar
DOI: 10.1101/gr.121905.111
发表时间: 2011-12-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Lee, Dongwon;Karchin, Rachel;Beer, Michael A.
通讯作者: Beer, Michael A.