Multivariate Nonparametric Methodology Studies
Multivariate Nonparametric Methodology Studies
批准号:
9626187
负责人:
David Scott
金额:
$0.0万
依托单位国家:
美国
项目类别:
Continuing grant
财政年份:
1996
资助国家:
美国
项目状态:
已结题
起止时间:
1996-06-15 至 1999-08-31
中文摘要
斯科特非参数方法学广泛应用于一维和二维,但不应用于高维。本研究侧重于中等维度,并对维度诅咒的含义和与大规模数据集相关的相关问题提供了更深入的理解。特别强调多元回归和密度估计问题,以及密切相关的应用,如聚类和脊。轶事证据表明,非参数方法在实践中的明显成功与理论预测的不良表现之间存在差距。我们研究了新的观点,特别是与局部自适应估计有关的观点。高质量的估计通常需要使用负核,但我们的结果表明,在黑森不确定的区域,通常在高维中占主导地位的尾部,等效增益是可能的。此外,我们已经开发了一类局部自适应但不是高阶的算法,它在实际问题中工作得更好,并避免了负性问题。我们已经用几种方法解决了由高维引起的问题。我们已经创建了从密度估计的角度寻找有趣子空间的算法。这样的子空间由最大偏置内容定义,依次剥离低偏置子空间。我们已经研究了密度估计的半参数模型,它可以比普通的非参数算法更好地工作,将可行性扩展了几个额外的维度。在处理中等维度数据和不断增长的海量数据集时,可视化尤为重要。一个新的可视化工具的例子是密度大巡展,它执行普通的大巡展,但显示导出的密度估计的实时视图,即平均偏移直方图。我们发现穿越脊线和等高线对于控制或约束观看是有用的。我们已经将我们的密度可视化能力扩展到回归曲面以及视觉聚类和视觉识别应用中的相关问题。可视化对于组织复杂的多测试问题聚类也很重要,例如我们在基于模式树的模式估计和测试中的结果。我们研究了一种局部测试坍塌模式的算法,作为改进聚类算法的基础。一个自然的扩展已经证明了多处理器和大规模数据集的并行架构。海量数据集给数学科学带来了巨大的挑战。在最近的一次国家研究委员会研讨会上,许多科学家确定了他们工作中的关键统计需求:主成分的替代方案,用于探索大量数据的专门可视化工具,更好的聚类算法,以及处理非平稳数据的技术。我们的研究结果直接影响到这四个关键机会中的三个。这个程序代表了对多元估计中重要数据分析问题的主机的全面和长期攻击。不需要明确写下公式的统计技术被称为非参数方法,其中包括众所周知的直方图作为一个简单的例子。这种技术广泛用于一维和二维数据,但不适用于高维数据,而高维数据是大多数重大挑战问题存在的地方。这项研究的重点是在许多严肃的理论家已经表达了对非参数方法可能不起作用的担忧的中程维度。然而,众所周知,许多实践科学家和工程师已经成功地使用非参数方法处理来自信号处理、图像理解、数据挖掘等各种实际问题的数据。这项研究提供了对所谓的维度诅咒的含义和与大量数据集相关的特定问题的更深层次的理解。特别强调在多元回归和密度估计问题,以及密切相关的应用,如聚类和脊。我们对局部自适应估计如何克服非参数方法在几个维度上的局限性有了新的认识。对于高维数据,我们开发了从密度估计的角度寻找最有趣子空间的算法。这些子空间由最大偏置内容定义并依次构造,剥离低偏置子空间。在二维之外,可视化是一项关键任务,特别是与不断增长的海量数据集相关的任务。一个成功的例子是我们的新密度大旅行,它提供了一种实时查看高维数据的新方法。我们已经将密度估计可视化功能扩展到回归曲面。可视化对于检查数据以检测集群的存在也非常有用。这些集群对于确定为诸如字符识别、遥感作物识别、地下水污染以及许多更专业的工程和科学应用等建议所收集的数据的有用性至关重要。这些算法的多处理器和并行架构版本在海量数据集的情况下特别相关。在数学科学中,处理大量数据集是一个巨大的挑战。在最近的一次国家研究委员会研讨会上,许多科学家确定了他们工作中的关键统计需求:主成分的替代方案,用于探索大量数据的专门可视化工具,更好的聚类算法,以及处理非平稳数据的技术。我们的研究结果直接影响到这四个关键机会中的三个。这个程序代表了对多元估计中重要数据分析问题的主机的全面和长期攻击。非参数方法学似乎在专家手中工作得很好,本研究的目的不仅是帮助专家,而且是为了促进更广泛的受众使用该方法学。* * *
英文摘要
DSM9616187 Scott Nonparametric methodology is widely used in one and two dimensions, but not in high dimensions. This research focuses on the mid-range dimensions and provides a deeper understanding of the implications of the curse of dimensionality and related problems associated with massive data sets. Particular emphasis has been given to multivariate regression and density estimation problems, and closely related applications such as clustering and ridges. Anecdotal evidence has suggested a gap between the apparent successes of nonparametric methodology in practice and the poor performance predicted by theory. We have examined new points of view, especially related to locally adaptive estimation. Higher quality estimation has often required use of negative kernels, but our results have shown that equivalent gains are possible in regions where the Hessian is indefinite, often in the tails which dominate in higher dimensions. In addition, we have developed a class of locally adaptive but not higher order algorithms that work better in practical problems and avoid problems of negativity. We have addressed problems arising from high dimensionality in several ways. We have created algorithms for finding interesting subspaces from the density estimation point of view. Such subspaces are defined by maximal bias content, sequentially peeling off low bias subspaces. We have examined semiparametric models for density estimation that can work better than ordinary nonparametric algorithms, extending feasibility by several extra dimensions. Visualization is especially important when dealing with medium dimensional data and the growing body of massive data sets. One example of a new visualization tool is provided by the density grand tour, which performs an ordinary grand tour but displays a real-time view of a derived density estimate, the averaged shifted histogram. We have found that traversing ridges and contours is useful to control or constrain viewing. We h ave extended our density visualization capabilities to regression surfaces and related problems in visual clustering and visual discrimination applications. Visualization is also important for organizing complicated multiple testing problems is clustering, such as our results in mode estimation and testing based on the mode tree. We have investigated a local testing algorithm for collapsing modes as the basis for an improved clustering algorithm. A natural extension has been demonstrated for multiprocessor and parallel architectures for massive data sets. A great challenge in mathematical sciences is provided by massive data sets. At a recent National Research Council workshop, numerous scientists identified critical statistical needs in their work: alternatives to principal components, specialized visualization tools for exploring massive data, better clustering algorithms, and techniques for handling nonstationary data. Results from our research directly impact three of these four critical opportunities. This program represents a comprehensive and long-term attack on a host of important data analytic problems in multivariate estimation. %%% Statistical techniques that do not require formulae to be written down explicitly are called nonparametric methods and include the well-known histogram as a simple example. Such techniques are widely used with data in one and two dimensions, but not in higher dimensions where most of the grand challenge problems are to be found. This research focuses on the mid-range dimensions where many serious theoreticians have expressed concern that nonparametric methods may not work. However, it is well-known that many practicing scientists and engineers have been successfully using nonparametric methods with data from signal processing, image understanding, data mining, among a wide array of real problems. This research is providing a deeper understanding of the implications of the so-called curse of dimensionality and particular p roblems associated with massive data sets. Particular emphasis is given to problems in multivariate regression and density estimation, as well as closely related applications such as clustering and ridges. We have obtained a new understanding of how locally adaptive estimation should work in overcoming the usual limitations of nonparametric methodology in several dimensions. For higher dimensional data, we have developed algorithms for finding maximally interesting subspaces from the density estimation point of view. Such subspaces are defined by maximal bias content and are constructed sequentially, peeling off low bias subspaces. Beyond two dimensions, visualization is a critical task, especially as related to the growing body of massive data sets. One example of a success is provided by our new density grand tour, which provides a new way of looking at high dimensional data in real-time. We have extended our density estimation visualization capabilities to regression surfaces. Visualization is also very useful for examining data to detect the presence of clusters. Such clusters are critical for determining the usefulness of data collected for proposes such as character recognition, remote sensing crop identification, ground water pollution, as well as many more specialized engineering and scientific applications. Multiprocessor and parallel architectures versions of these algorithms are particularly relevant in the massive data set situation. A great challenge in mathematical sciences is provided by handling massive data sets. At a recent National Research Council workshop, numerous scientists identified critical statistical needs in their work: alternatives to principal components, specialized visualization tools for exploring massive data, better clustering algorithms, and techniques for handling nonstationary data. Results from our research directly impact three of these four critical opportunities. This program represents a comprehensive and long-term attack on a host of important data analytic problems in multivariate estimation. Nonparametric methodology seems to work well in the hands of experts, and this research is designed to not only aid the expert but to facilitate the use of the methodology by a wider audience. ***
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Doctoral Dissertation Research: Comparing Multi-Scalar Claims for Redress and Reparation
-
批准号:1823901
-
项目类别:Standard Grant
-
资助金额:$2.52万
-
财政年份:2018
-
负责人:David Scott
-
依托单位:
17ALERT bid: A new multi-wavelength analytical ultracentrifuge for the study of biomolecular interactions
-
批准号:BB/R013411/1
-
项目类别:Research Grant
-
资助金额:$52.57万
-
财政年份:2018
-
负责人:David Scott
-
依托单位:
Multivariate Nonparametric Methodology Studies
-
批准号:0907491
-
项目类别:Standard Grant
-
资助金额:$10.0万
-
财政年份:2009
-
负责人:David Scott
-
依托单位:
Fluorescence Optics for the Analytical Ultracentrifuge
-
批准号:BB/F011156/1
-
项目类别:Research Grant
-
资助金额:$14.86万
-
财政年份:2008
-
负责人:David Scott
-
依托单位:
Multivariate Nonparametric Methodology Studies
-
批准号:0505584
-
项目类别:Continuing grant
-
资助金额:$0.0万
-
财政年份:2005
-
负责人:David Scott
-
依托单位:
Systemic Thread-Based Adaptation of an Electrical Engineering Curriculum
-
批准号:0343297
-
项目类别:Standard Grant
-
资助金额:$10.0万
-
财政年份:2003
-
负责人:David Scott
-
依托单位:
Multivariate Nonparametric Methodology Studies
-
批准号:0204723
-
项目类别:Continuing grant
-
资助金额:$0.0万
-
财政年份:2002
-
负责人:David Scott
-
依托单位:
Digital Government: Collaborative Research: Quality Graphics for Federal Statistical Summaries
-
批准号:9983459
-
项目类别:Continuing grant
-
资助金额:$27.0万
-
财政年份:2000
-
负责人:David Scott
-
依托单位:
Multivariate Nonparametric Methodology Studies
-
批准号:9971797
-
项目类别:Continuing grant
-
资助金额:$0.0万
-
财政年份:1999
-
负责人:David Scott
-
依托单位:
SBIR Phase I: Novel Inexpensive Titanium Dioxide-Assisted Photocatalysis for Waste Stream Remediation
-
批准号:9861306
-
项目类别:Standard Grant
-
资助金额:$9.98万
-
财政年份:1999
-
负责人:David Scott
-
依托单位:
Mathematical Sciences: Workshop on Advances in Smoothing: Bumps, Jumps, Clustering and Discrimination; May 11-15, 1997; Houston, Texas
-
批准号:9615912
-
项目类别:Standard Grant
-
资助金额:$0.0万
-
财政年份:1997
-
负责人:David Scott
-
依托单位:
Research Conference: Computing Science and Statistics Interface Symposium to be held May 14-17, 1997 in Houston, Texas
-
批准号:9708176
-
项目类别:Standard Grant
-
资助金额:$0.0万
-
财政年份:1997
-
负责人:David Scott
-
依托单位:
RUI: Genetic Analysis of Drosophila Pheromones
-
批准号:9614934
-
项目类别:Standard Grant
-
资助金额:$19.7万
-
财政年份:1997
-
负责人:David Scott
-
依托单位:
Support for Conference on Process Tomography
-
批准号:9619917
-
项目类别:Standard Grant
-
资助金额:$1.0万
-
财政年份:1997
-
负责人:David Scott
-
依托单位:
A Multidisciplinary Computer Integrated Freshman Level Circuits Laboratory with Practical Applications
-
批准号:9451961
-
项目类别:Standard Grant
-
资助金额:$6.54万
-
财政年份:1994
-
负责人:David Scott
-
依托单位:
Mathematical Sciences Computing Research Environments
-
批准号:9305700
-
项目类别:Standard Grant
-
资助金额:$2.08万
-
财政年份:1993
-
负责人:David Scott
-
依托单位:
1993 Presidential Awardees
-
批准号:9354288
-
项目类别:Standard Grant
-
资助金额:$0.75万
-
财政年份:1993
-
负责人:David Scott
-
依托单位:
Mathematical Sciences: Multivariate Nonparametric Methodology Studies
-
批准号:9306658
-
项目类别:Standard Grant
-
资助金额:$0.0万
-
财政年份:1993
-
负责人:David Scott
-
依托单位:
RUI: A Genetic Analysis of the Production and Perception ofPheromonal Signals
-
批准号:8906142
-
项目类别:Standard Grant
-
资助金额:$16.18万
-
财政年份:1989
-
负责人:David Scott
-
依托单位:
The Ultrastructural Correlates of Neural Transplantation and Development
-
批准号:8709687
-
项目类别:Standard Grant
-
资助金额:$2.0万
-
财政年份:1987
-
负责人:David Scott
-
依托单位:
海外基金