Robustly Identifying Dependent Components in Multiple High-Dimensional Data Sets Based on Few Observations
Robustly Identifying Dependent Components in Multiple High-Dimensional Data Sets Based on Few Observations
批准号:
262301625
负责人:
Professor Peter Schreier, Ph.D.
金额:
$0.0万
依托单位国家:
德国
项目类别:
Research Grants
财政年份:
2014
资助国家:
德国
项目状态:
已结题
起止时间:
2013-12-31 至 2021-12-31
中文摘要
这项提议的目标是开发方法,以便在样本支持相对较少的多个高维数据结构中稳健地识别依赖组件。许多算法需要此信息作为输入参数。例如,在生物医学中,已经建立了融合来自不同脑成像模式的数据的方法,但为了应用这些方法,我们需要知道不同特征集中的相关成分。作为另一个例子,在传感器阵列处理中,许多用于解析源的算法(例如,估计其到达方向)需要具有关于撞击阵列的源的数量的先验信息。通常,这个问题是临时解决的,结果大相径庭。因此,自然科学和工程学的许多领域都对系统方法的发展感兴趣。我们建议的重点将放在理论上,但为了说明和研究我们方法的性能,我们将选择一些在生物医学中的应用。更具体地说,在这项提议中,我们的目标是:-为样本支持相对较少的多个数据集制定模型选择规则。处理多个数据集比找到两个数据集之间的依赖关系要困难得多,因为存在许多可能的依赖关系结构。现有的极少数方法仅适用于大样本支持,并对潜在的相关性结构做出非常限制性的假设。-使我们的二阶技术对高斯性偏差具有健壮性。为了能够处理重尾噪声和异常值,这一点至关重要,这些噪声和异常值在许多应用中很常见。-首先建立一个基于二阶相关性的理论,该理论考虑数据集之间的线性相关性,然后扩展我们的方法,也考虑非线性相关性。-调查小样本支持对非线性依赖的可辨识性施加的限制。显然,我们预计样本的数量决定了可以从数据集中提取的信息量。-将我们的技术应用于生物医学中的一些精选问题。预计这些应用将从这一研究项目中受益匪浅。在生物医学界,使用特别的方法和经验法则来解决模型选择问题仍然很常见。系统的方法将有助于提供更令人信服和满意的解决方案。
英文摘要
The objective of this proposal is the development of methods to robustly identify dependent components in multiple high-dimensional data structures, where sample support is relatively small. Many algorithms require this information as an input parameter. For instance, in biomedicine, there are established approaches for fusing the data from different brain imaging modalities, but in order to apply them, we need to know the dependent components in different feature sets. As another example, in sensor array processing, many algorithms for resolving sources (e.g. estimating their direction of arrival) need to have prior information about the number of sources impinging upon the array. Often, this problem is solved ad hoc, with greatly varying results. The development of systematic approaches will therefore be of interest to a wide array of areas in the natural sciences and engineering. The focus of our proposal will be on the theory, but in order to illustrate and investigate the performance of our methods, we will choose some selected applications in biomedicine. More specifically, in this proposal, our objectives are: - To develop model-selection rules for multiple data sets with relatively small sample support. Treating multiple data sets is much more difficult than finding dependencies between two data sets because there are many possible dependence structures. The very few existing approaches work only for large sample support and make very restrictive assumptions about the underlying correlation structure.- To make our second-order techniques robust against deviations from Gaussianity. This is critical in order to be able to deal with heavy-tailed noise and outliers, which are commonplace in many applications. - To first build a theory based on second-order correlations, which consider linear dependencies between data sets, and then extend our approaches to also take into account nonlinear dependencies. - To investigate the restrictions that small sample support imposes on the identifiability of nonlinear dependencies. Obviously, we would expect that the number of samples determines the amount of information that can be extracted from the data sets. - To apply our techniques to some selected problems in biomedicine. It is expected that these applications will benefit greatly from this research project. It is still commonplace in the biomedical community to solve model-selection problems using ad-hoc approaches and rules of thumb. A systematic approach will help provide more convincing and satisfying solutions.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Nonparametric Techniques for Analyzing Directional Structure in Space-Time Random Fields
-
批准号:239765482
-
项目类别:Research Grants
-
资助金额:$0.0万
-
财政年份:2013
-
负责人:Professor Peter Schreier, Ph.D.
-
依托单位:
海外基金