课题基金 / 基金详情

dCortools: Distance Correlation Methods for Detecting Nonlinear Associations in High-Dimensional Molecular Data

dCortools: Distance Correlation Methods for Detecting Nonlinear Associations in High-Dimensional Molecular Data
dCortools:用于检测高维分子数据中非线性关联的距离相关方法
批准号:
417754611
负责人:
Dr. Dominic Edelmann
金额:
$0.0万
依托单位:
依托单位国家:
德国
项目类别:
Research Grants
财政年份:
2019
资助国家:
德国
项目状态:
已结题
起止时间:
2018-12-31 至 2021-12-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
目前用于测试高维分子数据关联的几乎所有方法都只能检测线性或单调关联。这既涉及不同分子变量之间的关联性测试(如基因-基因-相互作用),也涉及分子与临床变量之间的关联性测试(例如基因-环境-相互作用)。然而,众所周知,许多生物学关系更加复杂,包括非单调甚至非功能依赖。距离相关性是一种新的相关性度量,可以检测任意维随机向量之间的各种相关性。此外,距离相关系数很容易计算,这为它在统计实践中的应用奠定了基础。尽管有这些令人信服的性质,但到目前为止,距离相关系数在高维分子数据上的应用很少。这一方面是由于缺少生物统计学问题的方法学,另一方面是因为缺乏面向应用的软件。该项目的目标是缩小这一差距。在该项目的第一部分,我们计划开发用于生物医学应用的距离相关方法。首先,我们的目标是在强相关性结构的假设下,推导出比单变量程序更有效的迭代变量选择程序,这通常存在于分子数据中。此外,我们建议将距离相关系数扩展到生存数据,这在癌症研究中特别重要。对于项目的第二部分,我们计划创建一个用户友好的R包,它结合了对生物统计学有用的距离相关方法,从而允许从业者应用这种方法。在项目的第一部分中开发的技术将是这个R包的重要组成部分。最后,我们建议将R包应用于DACHES研究中的一个数据集,包括表观基因组范围的甲基化数据、2000多名结直肠癌患者的流行病学和临床数据。我们相信,计划中的项目将导致距离相关方法在生物统计实践中的使用大幅增加。对于分子数据,这将允许检测复杂的关联,如果使用线性程序,这些关联将被遗漏。这反过来可能会导致对生物过程的更好理解。
英文摘要
Virtually all methods that are currently used for testing associations in high-dimensional molecular data can only detect linear or monotone associations. This concerns both tests for the association between different molecular variables (e.g. gene-gene-interactions) and tests for the association between molecular and clinical variables (e.g. gene-environment-interactions).However, it is known that many biological relations are more complex, including nonmonotone or even nonfunctional dependencies. Distance correlation is a novel dependence measure that can detect every kind of dependence between random vectors of arbitrary dimensions. Moreover, the distance correlation coefficient is very easy to compute, which predestines it for the application in statistical practice. In spite of these convincing properties, there are hitherto only few applications of the distance correlation coefficient on high-dimensional molecular data. This is due to missing methodology for biostatistical problems on the one hand and to a lack of application-oriented software on the other hand. The goal of this project is to close this gap. In the first part of the project, we plan to develop distance correlation methodology for biomedical applications. First, we aim to derive iterative variable selection procedures that are much more efficient than univariate procedures under the assumption of strong correlation structures, which are typically present in molecular data. Moreover, we propose to extend the distance correlation coefficient to survival data, which are particularly important in cancer research.For the second part of the project we plan to create a user friendly R package that combines distance correlation methods that are useful for biostatistics and hence allows the application of this methodology for the practitioner. The techniques developed in the first part of the project will be important components of this R package. Finally, we propose to apply the R package on a data set from the DACHS study, consisting of epigenome-wide methylation data, epidemiological and clinical data for more than 2000 patients with colorectal cancer.We are confident that the planned project will lead to a considerable increase of the use of distance correlation methodology in biostatistical practice. For molecular data, this will allow to detect complex associations that would be missed if linear procedures were used. This in turn may lead to a better understanding of biological processes.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金