课题基金 / 基金详情

Symbolic Inference for Very Large Datasets

Symbolic Inference for Very Large Datasets
非常大的数据集的符号推理
批准号:
0805245
负责人:
Lynne Billard
金额:
$15.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2008
资助国家:
美国
项目状态:
已结题
起止时间:
2008-08-01 至 2012-07-31

项目摘要

项目成果

Lynne Billard的其他基金

相似基金

相关文献

中文摘要
翻译
由于数据本身“复杂”(和/或结构造成复杂性)而变得复杂的数据集,随着当代计算机能力的影响,正变得越来越常规。不寻常的是如何分析这些数据。事实上,数据“收集”的速度快于分析数据的能力。显然,即使在理论上似乎适用现有方法的情况下,常规使用这种统计技术往往是不适当的。有些方法(如压扁法)取有代表性的“样本”,然后对抽样数据使用标准程序。其他人则寻找子模式(例如,数据挖掘),然后尝试关注这些模式背后的数据。其他人则以某种有意义的方式汇总数据。其中一种聚合方法产生所谓的符号数据(如列表、区间、分布等)。符号数据的一个优点是,与采样集不同,符号值保留了所有原始数据,同时减少了数据集的大小。此外,虽然今天遇到的大量数据集是符号数据的一个来源,但有许多数据自然是符号的(无论是这些小数据集还是大数据集)。所有这些都可以用为符号数据开发的方法进行更好的分析。调查涉及三个主要领域。一个领域是分类树。在这里,开发了间隔和直方图值数据的距离度量;然后将它们用于新算法中,将经典的CART方法扩展到符号数据。其次,将回归方法,特别是逻辑回归和Cox比例风险模型,应用于符号数据。最后,对符号数据进行了因子分析和主成分分析。在当代计算机能力的影响下,数据本身“复杂”的数据集正变得越来越普遍。然而,这些计算机往往缺乏分析这些海量数据集的能力。因此,必须开发新的方法来处理它们。一种方法是以科学上有意义的方式聚合数据(实际的聚合由手头的问题决定)。这种聚合必然会产生以列表、间隔、直方图等形式出现的数据。研究者在三个主要领域开发了区间数据的新方法,区间和直方图的间隔距离测量后的分类树,回归方法,特别是逻辑回归和因子分析。结果应用于数据。通过整合数学/统计/计算领域来解决当代数据集遇到的实际问题,可以实现协同作用。这些成果不能仅靠其中一个学科的工具来实现,而是需要所有三个学科的工具。新的方法将广泛适用于气象学、环境科学、社会科学、卫生保健方案、工业等领域产生的数据集,远远超出了推动这项工作的那些数据集。这将对美国科学产生巨大影响。此外,由于博士生将作为合作者参与,国际研究人员将积极参与,这项研究有助于下一代和下一代美国科学家的国际化。
英文摘要
Datasets that are complex with the data themselves "complex", and/or with structures that impose complications) are becoming more and more routine with the impact of contemporary computer capacity. What is not routine is how to analyse these data. Indeed, the data "collection" is fast outpacing the ability to analyse them. It is evident that, even in those situations where in theory available methodology might seem to apply, routine use of such statistical techniques is often inappropriate. Some methods (e.g. squashing) take representative "`samples"' and then use standard procedures on the sampled data. Others seek sub/patterns (e.g., data mining) and then try to focus on the data behind those patterns. Others aggregate the data in some meaningful way. One such aggregation method produces so-called symbolic data (such as lists, intervals, distributions, etc.). An advantage of symbolic data is that unlike those in sampled sets, a symbolic-value retains all the original data, while simultaneously reducing the size of the dataset. Further while the massive datasets encountered today are one source of symbolic data, there are many data that are naturally symbolic (be these small or large datasets). All are better analysed by methods developed for symbolic data. The investigator addresses three major areas. One area is classification trees. Here, distances measures for interval and histogram-valued data are developed; and then they are used in new algorithms which extend the classical CART methodolgy to symbolic data. Secondly, regression methods, in particular, logistic regression and Cox's proportional hazard models, are adapted to symbolic data. Finally, factor analysis and principal component methodoly is developed for symbolic data. With the impact of contemporary computer capacity, datasets that are complex with the data themselves "complex" are becoming more ubiquitous. Yet those same computers often lack the capacity to analyse these massive datasets. Therefore, new ways to handle them must be developed. One way is to aggregate the data in a scientifically meaningful way (with the actual aggregation being dictated by the question at hand). Such aggregation will necessarily produce data that form lists, intervals, histograms, etc. The investigator develops new methodologies for interval data in three major areas, classification trees after rst nding distance measures for intervals and histograms, regression methods especially logistic regression, and factor analysis. The results are applied to data. A synergism is achieved by the integration of mathematical/ statistical/computational arenas in addressing real issues encountered by contemporary datasets. The outcomes cannot be achieved by the tools of just one of these disciplines but needs all three. The new methodologies will have wide applicability to those datasets generated in, e.g., meteorology, environmental science, social sciences, health-care programs, industry, and the like, well beyond those motivating the work. This will have enormous impact on US science. Further since doctoral students will be engaged as collaborators and since international researchers will be active participants, the research helps in the internationalization of the next and future generation of US scientists.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Workshop: Pathways to the Future Workshop 2004
Statistical Inference for Complex Data
Pathways to the Future Workshop 2003
U.S.-France Cooperative Research (INRIA): Symbolic Data Analysis Project
海外基金