Two-Tier Mapper, an unbiased topology-based clustering method for enhanced global gene expression analysis

Two-Tier Mapper, an unbiased topology-based clustering method for enhanced global gene expression analysis
复制标题

DOI:
10.1093/bioinformatics/btz052
复制
发表时间:
2019-09-15
期刊:
影响因子:
5.8
通讯作者:
Brisken, Cathrin
Brisken, Cathrin
中科院分区:
生物学3区
文献类型:
--
作者:
Jeitziner, Rachel;Carriere, Mathieu;Brisken, Cathrin

文献摘要

被引文献

相似文献

动机:需要无偏聚类方法来分析越来越多的复杂数据集。现有的聚类方法往往依赖于用户设置的参数,缺乏稳定性,不适用于小数据集。为了克服这些缺点,我们使用拓扑数据分析,一个新兴的数学领域,辨别额外的功能,发现隐藏的见解datasets和具有广泛的应用范围。结果:我们已经开发了一种基于拓扑的聚类方法称为两层映射(TTMap),用于增强全球基因表达数据集的分析。首先,TTMap辨别出对照组中的不同特征,对其进行调整,并识别出离群值。其次,每个测试样本从控制组在高维空间中的偏差计算,和测试样本聚类使用一种新的基于映射器的拓扑算法在两个层次上:一个全局层和局部层。所有参数都经过精心选择或数据驱动,避免任何用户引起的偏差。该方法是稳定的,不同的数据集可以组合进行分析,并可以识别显着的亚组。它在合成和生物数据集的灵敏度和稳定性方面优于当前的聚类方法,特别是当样本量很小时;结果不受去除对照样本、选择归一化或选择数据的影响。TTMap很容易应用于复杂的,高度可变的生物样品,并为个性化医疗带来希望。
Motivation: Unbiased clustering methods are needed to analyze growing numbers of complex datasets. Currently available clustering methods often depend on parameters that are set by the user, they lack stability, and are not applicable to small datasets. To overcome these shortcomings we used topological data analysis, an emerging field of mathematics that discerns additional feature and discovers hidden insights on datasets and has a wide application range.Results: We have developed a topology-based clustering method called Two-Tier Mapper (TTMap) for enhanced analysis of global gene expression datasets. First, TTMap discerns divergent features in the control group, adjusts for them, and identifies outliers. Second, the deviation of each test sample from the control group in a high-dimensional space is computed, and the test samples are clustered using a new Mapper-based topological algorithm at two levels: a global tier and local tiers. All parameters are either carefully chosen or data-driven, avoiding any user-induced bias. The method is stable, different datasets can be combined for analysis, and significant subgroups can be identified. It outperforms current clustering methods in sensitivity and stability on synthetic and biological datasets, in particular when sample sizes are small; outcome is not affected by removal of control samples, by choice of normalization, or by subselection of data. TTMap is readily applicable to complex, highly variable biological samples and holds promise for personalized medicine.