课题基金 / 基金详情

Collaborative Research: Information Matrix Analysis for Nonparametric Multivariate Problems

Collaborative Research: Information Matrix Analysis for Nonparametric Multivariate Problems
协作研究:非参数多元问题的信息矩阵分析
批准号:
1407639
负责人:
Francesca Chiaromonte
金额:
$24.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2014
资助国家:
美国
项目状态:
已结题
起止时间:
2014-08-15 至 2019-07-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
该项目开发了一套新的工具,称为信息矩阵分析(IMA),以探索多变量数据的结构。有了现代数据采集设备和广阔的数据存储空间,研究人员可以轻松地收集高维数据,如生物技术数据、金融数据、卫星图像和高光谱图像。这种高维数据的分析给统计学家带来了巨大的挑战,这是由于所谓的“维度诅咒”。IMA方法通过找到较少数量的原始变量的线性组合来解决这个问题,这些组合将携带原始多变量数据的大部分信息。IMA基于针对不同问题定义的信息矩阵的特征分析。所提出的信息矩阵的特征向量提供了变量的线性组合,这些变量最好地总结了数据的有用特征,因为它们的信息含量很高。通过将新方法与随机投影法相结合,还可以将IMA应用于超高维数据。或者,可以使用IMA进行变量选择。由于数据革命,该项目正在开发的新数据分析工具是及时的,而且具有广泛的适用性。所有的IMA方法都源于一个共同的基础,因此很容易移植到解决新问题的时候。该项目将大大加强统计工具和软件的可用性,以便对多变量数据进行统计建模和探索。新的方法将使想要分析包括医学研究、预防研究、公共卫生和社会科学在内的各个领域的高维数据的科学家和研究人员受益。该项目将加速对菲舍尔信息矩阵的几何理解。该项目的核心是将IMA应用于三个非常不同的问题领域:建立模型、评估模型和比较人口。假设有人想要建立一个模型来分析响应变量和多变量协变量之间的关系。定义的协变量信息矩阵的IMA可以用来寻找原始协变量的较少数量的线性投影,以简化模型的建立。新的投影变量很好地解释了,通过定义的协变量信息来衡量,携带了关于协变量和响应之间的关系的大多数信息。IMA还可以推广到找到数据的线性组合,从而在两个密度之间提供最佳的区分。在应用中,一个密度是真实的未知密度,非参数估计,另一个密度是数据的某种模型,可以是参数的或半参数的。或者,这两种密度可以代表人们希望比较的两个不同的种群。当两个总体是具有相同协方差矩阵的多元正态分布时,IMA提供与Fisher线性判别方向相同的线性投影。然而,IMA可以超越线性判别分析,转向多个线性判别分析。IMA将被进一步应用于寻找数据的线性投影,这些线性投影对于评估所提议的模型的适合性是有用的,无论是参数模型、半参数模型还是非参数模型。本项目将研究IMA在流行的图形模型和独立成分分析模型中的应用。预计IMA可能会有更多的应用领域。对于高维数据,该建议将发展使用随机投影方法,首先降低原始变量的维度,并限制可能的线性组合的空间。然后,IMA可以应用于低得多的维数据集。或者,人们可以通过施加信息投影的稀疏性来使用IMA进行变量选择。
英文摘要
This project develops a new set of tools, called Information Matrix Analysis (IMA), to explore the structure of multivariate data. With modern data gathering devices and vast data storage space, researchers can easily collect high-dimensional data, such as biotech data, financial data, satellite imagery, and hyperspectral imagery. Analysis of such high-dimensional data poses great challenges for statisticians due to the so-called "curse of dimensionality". The IMA methodology tackles this problem by finding a smaller number of linear combinations of the original variables that will carry most of the information of the original multivariate data. The IMA is based on the eigenanalysis of the Information Matrix as defined for different problems. The eigenvectors of the proposed information matrices provide the linear combinations of the variables that best summarize the useful features of the data due to their high information content. By combining the new method with the random projection method, one can also apply IMA to ultra high dimensional data. Alternatively, one can do variable selection with IMA. The new data analysis tool under development in this project is timely due to the data revolution, as well as broadly applicable. All of the IMA methodology springs from a common foundation, and so is easily transported to tackle new problems. This project will enhance significantly the availability of statistical tools and software for statistical modeling and exploration for multivariate data. The new method will benefit a broad range of scientists and researchers who want to analyze high-dimensional data in various fields, including medical studies, prevention studies, public health, and the social sciences.This project will accelerate geometric understanding of Fisherian information matrices. The core of this project consists of the application of IMA to three very different problem areas: building models, assessing models, and comparing populations. Suppose one wants to build a model to analyze the relationship between a response variable and multivariate covariates. The IMA of a defined covariate information matrix can be applied to find a smaller number of linear projections of the original covariates to simplify model building. The new projected variables have the nice explanation of carrying most information, as measured by the defined covariate information, about the relationship between the covariates and the response. The IMA can be also generalized to find linear combinations of the data that can provide the best discrimination between two densities. In the applications, one density will be the true unknown density, to be estimated nonparametrically, and the other density will be some model for the data, which could be parametric or semiparametric. Alternatively, the two densities could represent two distinct populations one wishes to compare. When the two populations are multivariate normal with the same covariance matrix, the IMA provides the same linear projection as the Fisher linear discriminant direction. However, the IMA can move beyond linear discriminant analysis to multiple linear discriminants. The IMA will be further applied to find linear projections of the data that are useful for assessing the fit of a proposed model, whether parametric, semiparametric, or nonparametric. This project will investigate the applications of IMA to popular graphical models and independent component analysis models. It is expected that IMA might have many more application areas. For high dimensional data, the proposal will develop the use of the random projection method to first reduce the dimension of the original variables and restrict the space of possible linear combinations. Then the IMA can be applied to a much lower dimensional data set. Alternatively, one can do variable selection with IMA by imposing the sparsity of the informative projections.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
Research on Quantum Field Theory without a Lagrangian Description
  • 批准号:
    24ZR1403900
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    SATOSHI NAWATA
  • 依托单位:
Cell Research
Cell Research
Cell Research (细胞研究)