Omada: Robust clustering of transcriptomes through multiple testing

Omada: Robust clustering of transcriptomes through multiple testing
复制标题

DOI:
10.1101/2022.12.19.519427
复制
发表时间:
2022-12
期刊:
bioRxiv
影响因子:
--
通讯作者:
S. Kariotis;Tan Pei Fang;Haiping Lu;Christopher J. Rhodes;Martin Wilkins;A. Lawrie;Dennis Wang
S. Kariotis;Tan Pei Fang;Haiping Lu;Christopher J. Rhodes;Martin Wilkins;A. Lawrie;Dennis Wang
中科院分区:
其他
文献类型:
--
作者:
S. Kariotis;Tan Pei Fang;Haiping Lu;Christopher J. Rhodes;Martin Wilkins;A. Lawrie;Dennis Wang

文献摘要

相似文献

队列研究越来越多地收集生物样本进行分子分析,并观察分子异质性。高通量RNA测序提供了能够反映疾病机制的大型数据集。聚类方法已经产生了许多工具来帮助剖析复杂的异构数据集,然而,选择合适的方法和参数来对转录组学数据进行探索性聚类分析需要对机器学习和广泛的计算实验有深刻的理解。在没有事先的领域知识的情况下,帮助做出此类决策的工具是不存在的。为了解决这个问题,我们开发了Omada,一套旨在自动化这些过程的工具,并通过基于自动化机器学习的功能使转录组数据的鲁棒无监督聚类更容易获得。每个工具的效率用五个数据集进行测试,这些数据集具有不同的表达信号强度,以捕获广泛的RNA表达数据集。我们的工具包的决策反映了数据集中稳定分区的实际数量,其中子组是可识别的。在生物学差异不太明确的数据集中,我们的工具要么形成具有不同表达谱和强大临床关联的稳定亚组,要么揭示有问题数据的迹象,如有偏差的测量。
Cohort studies increasingly collect biosamples for molecular profiling and are observing molecular heterogeneity. High throughput RNA sequencing is providing large datasets capable of reflecting disease mechanisms. Clustering approaches have produced a number of tools to help dissect complex heterogeneous datasets, however, selecting the appropriate method and parameters to perform exploratory clustering analysis of transcriptomic data requires deep understanding of machine learning and extensive computational experimentation. Tools that assist with such decisions without prior field knowledge are nonexistent. To address this we have developed Omada, a suite of tools aiming to automate these processes and make robust unsupervised clustering of transcriptomic data more accessible through automated machine learning based functions. The efficiency of each tool was tested with five datasets characterised by different expression signal strengths to capture a wide spectrum of RNA expression datasets. Our toolkit’s decisions reflected the real number of stable partitions in datasets where the subgroups are discernible. Within datasets with less clear biological distinctions, our tools either formed stable subgroups with different expression profiles and robust clinical associations or revealed signs of problematic data such as biased measurements.