16S rRNA sequence embeddings: Meaningful numeric feature representations of nucleotide sequences that are convenient for downstream analyses

16S rRNA sequence embeddings: Meaningful numeric feature representations of nucleotide sequences that are convenient for downstream analyses
复制标题

DOI:
10.1371/journal.pcbi.1006721
复制
发表时间:
2018-05
影响因子:
4.3
通讯作者:
Stephen Woloszynek;Zhengqiao Zhao;Jian Chen;G. Rosen
Stephen Woloszynek;Zhengqiao Zhao;Jian Chen;G. Rosen
中科院分区:
生物学2区
文献类型:
--
作者:
Stephen Woloszynek;Zhengqiao Zhao;Jian Chen;G. Rosen

文献摘要

被引文献

相似文献

高通量测序的进步增加了微生物组测序数据的可用性,这些数据可以用来原位表征微生物组群落结构。我们探索了对核苷酸序列使用单词和句子嵌入方法,因为它们可能是下游机器学习应用(特别是深度学习)的合适数值表示。这项工作包括首先将每个序列编码(“嵌入”)到一个密集的、低维的数字向量空间中。在这里,我们使用Skip-Gram word2vec嵌入从16S rRNA扩增子调查中获得的k-mers,然后利用现有的句子嵌入技术嵌入属于特定身体部位或样本的所有序列。我们证明这些表征具有生物学意义,因此嵌入空间可以作为一种特征提取形式用于探索性分析。我们发现序列嵌入保留了有关测序数据的相关信息,如k-mer上下文、序列分类和样本类别。具体而言,序列嵌入空间解决了门之间的差异,以及同一科内属之间的差异。序列嵌入之间的距离与对齐身份之间的距离具有相似的性质,并且嵌入多个序列可以被认为是生成一致序列。与使用OTU丰度数据相比,使用样本嵌入进行身体部位分类导致的性能损失可以忽略不计。最后,k-mer嵌入空间捕获了不同的k-mer谱,这些谱映射到16S rRNA基因的特定区域,并与特定的身体部位相对应。总之,我们的研究结果表明,嵌入序列产生有意义的表示,可用于探索性分析或需要数字数据的下游机器学习应用。此外,由于嵌入以无监督的方式进行训练,因此可以嵌入未标记的数据并用于支持有监督的机器学习任务。基因组测序方式的改进导致了微生物组数据的丰富。通过正确的方法,研究人员使用这些数据来彻底表征微生物如何相互作用以及它们的宿主,但是测序数据的形式(字母序列)对于许多数据分析方法来说并不理想。因此,我们提出了一种将测序数据转换为数字数组的方法,可以在子序列、全序列和样本水平上捕获数据的有趣质量。这使我们能够衡量某些微生物序列对微生物类型和宿主条件的重要性。此外,以这种方式表示序列可以提高我们使用其他复杂建模方法的能力。使用来自人体样本的微生物组数据,我们表明我们的数字表示捕获了不同类型微生物之间的差异,以及收集样本的身体部位位置的差异。
Advances in high-throughput sequencing have increased the availability of microbiome sequencing data that can be exploited to characterize microbiome community structure in situ. We explore using word and sentence embedding approaches for nucleotide sequences since they may be a suitable numerical representation for downstream machine learning applications (especially deep learning). This work involves first encoding (“embedding”) each sequence into a dense, low-dimensional, numeric vector space. Here, we use Skip-Gram word2vec to embed k-mers, obtained from 16S rRNA amplicon surveys, and then leverage an existing sentence embedding technique to embed all sequences belonging to specific body sites or samples. We demonstrate that these representations are biologically meaningful, and hence the embedding space can be exploited as a form of feature extraction for exploratory analysis. We show that sequence embeddings preserve relevant information about the sequencing data such as k-mer context, sequence taxonomy, and sample class. Specifically, the sequence embedding space resolved differences among phyla, as well as differences among genera within the same family. Distances between sequence embeddings had similar qualities to distances between alignment identities, and embedding multiple sequences can be thought of as generating a consensus sequence. Using sample embeddings for body site classification resulted in negligible performance loss compared to using OTU abundance data. Lastly, the k-mer embedding space captured distinct k-mer profiles that mapped to specific regions of the 16S rRNA gene and corresponded with particular body sites. Together, our results show that embedding sequences results in meaningful representations that can be used for exploratory analyses or for downstream machine learning applications that require numeric data. Moreover, because the embeddings are trained in an unsupervised manner, unlabeled data can be embedded and used to bolster supervised machine learning tasks. Author summary Improvements in the way genomes are sequenced have led to an abundance of microbiome data. With the right approaches, researchers use this data to thoroughly characterize how microbes interact with each other and their host, but sequencing data is of a form (sequences of letters) not ideal for many data analysis approaches. We therefore present an approach to transform sequencing data into arrays of numbers that can capture interesting qualities of the data at the sub-sequence, full-sequence, and sample levels. This allows us to measure the importance of certain microbial sequences with respect to the type of microbe and the condition of the host. Also, representing sequences in this way improves our ability to use other complicated modeling approaches. Using microbiome data from human samples, we show that our numeric representations captured differences between different types of microbes, as well as differences in the body site location from which the samples were collected.