Discovering and deciphering relationships across disparate data modalities

Discovering and deciphering relationships across disparate data modalities
复制标题

DOI:
10.7554/elife.41690
复制
发表时间:
2019-01-15
期刊:
影响因子:
7.7
通讯作者:
Shen, Cencheng
Shen, Cencheng
中科院分区:
生物学1区
文献类型:
--
作者:
Vogelstein, Joshua T.;Bridgeford, Eric W.;Shen, Cencheng

文献摘要

被引文献

相似文献

了解数据的不同属性之间的关系,例如基因组或连接体是否具有关于疾病状态的信息,变得越来越重要。虽然现有的方法可以测试两个属性是否相关,但它们可能需要不可行的大样本量,并且通常无法解释。我们的方法,“多尺度图相关性”(MGC),是一个依赖性测试,并列不同的数据科学技术,包括k-最近邻,内核方法和多尺度分析。其他方法可能需要两倍或三倍的样本数量,以实现与基准套件中的MGC相同的统计功效,包括高维和非线性关系,维度范围从1到1000。此外,MGC唯一地表征潜在的几何关系,同时保持计算效率。在真实的数据中,包括脑成像和癌症遗传学,MGC检测依赖性的存在,并为下一步的实验提供指导。
Understanding the relationships between different properties of data, such as whether a genome or connectome has information about disease status, is increasingly important. While existing approaches can test whether two properties are related, they may require unfeasibly large sample sizes and often are not interpretable. Our approach, 'Multiscale Graph Correlation' (MGC), is a dependence test that juxtaposes disparate data science techniques, including k-nearest neighbors, kernel methods, and multiscale analysis. Other methods may require double or triple the number of samples to achieve the same statistical power as MGC in a benchmark suite including high-dimensional and nonlinear relationships, with dimensionality ranging from 1 to 1000. Moreover, MGC uniquely characterizes the latent geometry underlying the relationship, while maintaining computational efficiency. In real data, including brain imaging and cancer genetics, MGC detects the presence of a dependency and provides guidance for the next experiments to conduct.