A demonstration of unsupervised machine learning in species delimitation

A demonstration of unsupervised machine learning in species delimitation
复制标题

DOI:
10.1016/j.ympev.2019.106562
复制
发表时间:
2019-10-01
影响因子:
4.1
通讯作者:
Hedin, Marshal
Hedin, Marshal
中科院分区:
生物学1区
文献类型:
--
作者:
Derkarabetian, Shahan;Castillo, Stephanie;Hedin, Marshal

文献摘要

被引文献

相似文献

用遗传数据界定物种的一个主要挑战是成功区分种群结构和物种水平的差异,这一问题在栖息于自然破碎栖息地的类群中加剧。许多科学领域现在都在使用机器学习,在进化生物学中,监督机器学习最近被用来推断物种边界。这些监督方法需要具有相关标签的训练数据。相反,无监督机器学习(UML)使用固有的数据结构,不需要用户指定的训练标签,从而可能在物种划分方面提供更多的客观性。在综合分类学的背景下,我们展示了三个UML方法(随机森林,变分自动编码器,t-分布随机邻居嵌入)的实用程序,在一个蛛形纲动物类群与高人口遗传结构(Opiliones,Laniatores,Metanonychus)的物种界定。我们发现,UML方法成功地聚类样本,根据物种水平的分歧,而不是高层次的人口结构,而基于模型的验证方法严重过度分裂假定的物种。UML在二维空间中提供直观的数据可视化,能够适应各种数据类型,并且在系统和进化生物学的许多领域都有潜力。我们认为,机器学习方法非常适合物种划界,并可能在许多自然系统和具有不同生物学特征的类群中表现良好。
One major challenge to delimiting species with genetic data is successfully differentiating population structure from species-level divergence, an issue exacerbated in taxa inhabiting naturally fragmented habitats. Many fields of science are now using machine learning, and in evolutionary biology supervised machine learning has recently been used to infer species boundaries. These supervised methods require training data with associated labels. Conversely, unsupervised machine learning (UML) uses inherent data structure and does not require user-specified training labels, potentially providing more objectivity in species delimitation. In the context of integrative taxonomy, we demonstrate the utility of three UML approaches (random forests, variational autoencoders, t-distributed stochastic neighbor embedding) for species delimitation in an arachnid taxon with high population genetic structure (Opiliones, Laniatores, Metanonychus). We find that UML approaches successfully cluster samples according to species-level divergences and not high levels of population structure, while model-based validation methods severely over-split putative species. UML offers intuitive data visualization in two-dimensional space, the ability to accommodate various data types, and has potential in many areas of systematic and evolutionary biology. We argue that machine learning methods are ideally suited for species delimitation and may perform well in many natural systems and across taxa with diverse biological characteristics.