课题基金 / 基金详情

Matrix variate modeling and analysis of large scale biological data with complex dependencies

Matrix variate modeling and analysis of large scale biological data with complex dependencies
具有复杂依赖性的大规模生物数据的矩阵变量建模和分析
批准号:
1316731
负责人:
Kerby Shedden
金额:
$22.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2013
资助国家:
美国
项目状态:
已结题
起止时间:
2013-08-01 至 2017-07-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
矩阵变量模型提供了一种分析复杂多变量数据集的方法,其中变量之间和观察单位之间可能存在有意义的关系。 这是多变量统计分析的更标准设置的扩展,其中变量是相关的,但观察结果被视为独立的。 将数据视为单个高维样本的主要动机是在估计对实验单元之间的关系敏感的参数时增加功效的潜力。 高维非渐近理论、凸分析和算法的最新进展使这些模型能够适用于基因组学和神经科学等关键科学领域中出现的非常大的数据集。 该项目的一个主要目标是评估在何种程度上解释样本之间的关系,允许更准确地估计变量之间的关系。 这可能如何应用的一个例子是评估罕见遗传变异和表型之间的关联。 序贯核关联检验(SKAT)使用的核矩阵本质上是遗传特征的协方差矩阵。 研究者将评估从最近提出的Gemini矩阵变量分析方法中获得的协方差矩阵作为SKAT程序的核心的使用。三个科学应用已经确定,双子座框架可能会有益地扩展。 除了概念和算法挑战之外,这些应用还需要Gemini估计器的理论和方法来覆盖新的设置,例如,由于SNP基因型高度非高斯,以及由于大脑连接图随时间变化。 研究人员将使基线Gemini模型和算法适应这些新的设置,并研究它们的统计和计算特性。 生物学和健康科学的研究人员已经接受了这些技术,但有时会因为必须保守地解释这些数据而感到沮丧,以防止做出不可复制的声明。 如果有可能以一种提供更大统计效力的方式来分析这些数据,而不需要昂贵地增加样本量,那么研究的进展将会加快。 做到这一点的一个途径是超越独立处理研究中观察到的单位(例如人类研究对象或实验室动物)的范式。 研究人员和他的同事们开发了一种新型的统计程序,使用观察单位之间的推断关系来更准确地估计物理和生物未知数。 他们计划开发他们的技术背后的理论,开发用于进行分析的软件,并与癌症生物学,基因组学和神经科学的科学家合作,评估使用这种方法获得新的科学见解的潜力。 该项目将重点关注三个科学问题:识别与癌症发病率和进展有关的DNA病变,识别与人类疾病相关的遗传变异,以及表征由意识状态变化引起的神经连接变化。
英文摘要
Matrix-variate models provide a way to analyze complex multivariate datasets in which meaningful relationships may exist among both the variables and among the observed units. This is an extension of the more standard setting of multivariate statistical analysis, in which the variables are dependent, but the observations are viewed as being independent. The primary motivation for treating the data as a single high-dimensional sample is the potential for increased power when estimating parameters that are sensitive to relationships among the experimental units. Recent advances in high-dimensional non-asymptotic theory, convex analysis, and algorithms allow such models to be fit to the very large data sets that arise in critical scientific areas such as genomics and neuroscience. A major goal of this project is to assess the extent to which accounting for relationships among samples allows more accurate estimates to be made of relationships among variables. An example of how this might be applied is in the assessment of associations between rare genetic variants and a phenotype. The Sequential Kernel Association Test (SKAT) uses a kernel matrix which is essentially a covariance matrix of the genetic features. The investigators will evaluate the use of the covariance matrix obtained from the recently-proposed Gemini approach to matrix-variate analysis as a kernel for the SKAT procedure. Three scientific applications have been identified to which the Gemini framework may be usefully extended. In addition to conceptual and algorithmic challenges, these applications require the theory and methodology for the Gemini estimators to cover new settings, for example, due to SNP genotypes being highly non-Gaussian, and due to brain connectivity graphs changing over time. The investigators will adapt the baseline Gemini models and algorithms to these new settings, and study their statistical and computational properties.Recent technological breakthroughs in instrumentation allow large and detailed data sets describing living systems to be efficiently collected. Researchers in biology and health science have embraced these technologies, but have sometimes been frustrated by the fact that such data must be interpreted conservatively, to guard against making non-reproducible claims. Research progress would be accelerated if it were possible to analyze such data in a way that provides more statistical power, without expensive increases in the sample size. One path to doing this is to move beyond the paradigm of treating the observed units in a study (e.g. human research subjects or laboratory animals) independently. The investigator and his colleagues have developed a new type of statistical procedure that uses inferred relationships among the units of observation to more accurately estimate physical and biological unknowns. They plan to develop the theory behind their technique, to develop software for carrying out the analyses, and to work with scientists in cancer biology, genomics, and neuroscience to assess the potential for using this approach to obtain novel scientific insights. The project will focus on three scientific problems: identification of DNA lesions involved with cancer incidence and progression, identification of inherited genetic variants associated with human diseases, and characterization of neural connectivity changes induced by changes in consciousness state.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金