课题基金 / 基金详情

Statistical Inference for Complex Data

Statistical Inference for Complex Data
复杂数据的统计推断
批准号:
0400584
负责人:
Lynne Billard
金额:
$0.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2004
资助国家:
美国
项目状态:
已结题
起止时间:
2004-05-15 至 2009-04-30

项目摘要

项目成果

Lynne Billard的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
The advent of computers is routinely bringing large, very large datasets for analysis. How to analyze theattendant data and/or how to glean useful information from these data is anything but routine. This proposal investigates two broad areas, viz., mixture decomposition, and regression in the presence of taxonomy and hierarchical variables. The mixture decomposition work is concerned with data in which the observation is a distribution function in the p-dimensional Cartesian product of distributions space, rather than the single point in p-dimensional space of classical data. Such data arise naturally, or after aggregation of original data to a more manageable size yet retaining its inherent information. A goal is to partition these distributions into coherent classes and to estimate the relevant class distributions. It is proposed to adapt ideas from copula theory for classical data, to data comprised of distributions. Embedded in this process is the need to study parameter estimation. Different partitioning techniques will be explored such as dynamical clustering, different measures of fit criterion will be considered, e.g., log-likelihood classification; and different estimation methods explored for the underlying copulas and the associated distributions and their parameters, e.g., maximum likelihood, nonparametric methods such as Parzen's truncated window. The resulting methodology will have wide applicability to those datasets generated in, e.g., meteorology, environmental science, social sciences, health-care programs, and the like. Regression methods when taxonomy variables, and when hierarchical variables, are present will also be developed. This will first be executed for classical data, and then extended to interval-valued data and to histogram- (or frequency-) valued data.With modern computers generating very large datasets, it is imperativethat techniques be developed for analysing such datasets. To date veryfew methods exist. The research will develop new methodologies foranalysing these data. A first step is to aggregate the data in somewell-defined but meaningful way. This aggregation will thence producedata in the form of lists, intervals, or distributions, and these datawill now have some form of internal structure. The research will focuson such data of two types. One will be where the data now consist ofdistributions, and where the goal is to develop methods to identify theappropraite mixture of distributions that describe these data. Anotherdeals with taxonomy and hierarical data, with the goal of establishingregression relationships that will explain the underlying processgoverning the variables involved. The resulting methodologies willallow analysis and interpretation of contemporary datastes wherecurrently no methods exist.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Symbolic Inference for Very Large Datasets
Workshop: Pathways to the Future Workshop 2004
Pathways to the Future Workshop 2003
U.S.-France Cooperative Research (INRIA): Symbolic Data Analysis Project
海外基金