课题基金 / 基金详情

Statistical Theory and Methods for D&R Analysis of Large Complex Data

Statistical Theory and Methods for D&R Analysis of Large Complex Data
D 统计理论与方法
批准号:
1228348
负责人:
William Cleveland
金额:
$31.5万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2012
资助国家:
美国
项目状态:
已结题
起止时间:
2012-09-01 至 2017-08-31

项目摘要

项目成果

William Cleveland的其他基金

相似基金

相关文献

中文摘要
翻译
在分割和重组(D&R)中,数据分析师将数据分成子集。这些是S计算,因为它们创建了子集。统计和可视化方法应用于每个子集,计算之间没有通信。这些是W计算,因为它们在子集内。然后跨子集对W计算输出进行重组。这些是B计算,因为它们在子集之间。D&R的目标之一是深入分析,即无论数据大小和复杂程度如何,都能详细研究数据的能力。第二个目标是能够完全从交互式数据分析语言(ILDA)(如R)中执行分析。D&R通过引入简单的并行化来实现目标,而不是分析方法本身,这是非常复杂的,而是数据。这导致了“令人尴尬的并行”计算,可以由像Hadoop这样的分布式计算环境有效地执行。此外,Hadoop可以与ILDA合并。研究人员将研究两个领域的统计理论和方法的研发。首先是研发统计划分和重组程序的开发。这是非常广泛的,因为有许多分析方法,并且过程需要随着它们处理的方法和数据结构的变化而变化。第二个主题是基础数学理论。在当前的统计基本范式中,分析方法直接应用于一次大计算中的所有数据。S、W和B计算也使用所有数据,但结果通常与直接计算的结果不同,并且具有不同的统计属性。这为统计准确性和最优性引入了一个新的基本范式。在Divide and Recombine (D&R)中,将大型复杂数据划分为子集。统计和可视化方法分别应用于每个子集。然后将每种方法的结果跨子集进行重组。这种新的大型复杂数据分析框架可以很容易地利用当前的分布式计算环境,因为它导致非常简单的并行计算。研究人员将开发用于分裂和重组的统计程序,从而使分析方法具有良好的统计准确性。与在一次大计算中直接计算所有数据相比,这种方法的准确性往往更低,因为这种方法不切实际,而且时间长,或者根本不可行。D&R为了计算的可行性而牺牲了一些准确性。结果是,几乎任何统计或可视化方法都可以成功地应用于大型复杂数据。这样可以进行深入、详细的分析,而不会有丢失数据中重要信息的风险,这在今天只有小数据才可行。
英文摘要
In Divide and Recombine (D&R) the data are divided into subsets by the data analyst. These are the S computations because they create the subsets. Statistical and visualization methods are applied to each subset without communication among the computations. These are the W computations because they are within subsets. Then the W computation outputs are recombined across subsets. These are the B computations because they are between subsets. One goal of D&R is deep analysis, an ability to study the data in detail despite the size and complexity. A second goal is an ability to carry out analysis wholly from within an interactive language for data analysis (ILDA) such as R. D&R achieves the goals by introducing a simple parallelization, not of the analysis methods themselves which is very complex, but of the data. This results in ``embarrassingly parallel'' computations that can be efficiently carried out by a distributed computational environment like Hadoop. Also, Hadoop can be merged with an ILDA. The investigators will research two areas of statistical theory and methods for D&R. The first is development of D&R statistical division and recombination procedures. This is very broad because there are many analysis methods, and the procedures need to change with the methods and the data structures they address. The second topic is a foundational mathematical theory. In the current fundamental paradigm for statistics, an analysis method is applied directly to all of the data in one big computation. The S, W, and B computations use all of the data too, but the results are in general not the same as those for direct computation and have different statistical properties. This introduces a new fundamental paradigm for statistical accuracy and optimality.In Divide and Recombine (D&R), large complex data are divided into subsets. Statistical and visualization methods are applied to each of the subsets separately. Then the results of each method are recombined across subsets. This new analysis framework for large complex data can readily exploit current distributed computational environments because it leads to very simple parallel computation. The investigators will develop statistical procedures for division and recombination that result in good statistical accuracy for the analysis methods. Accuracy tends to be less than that from direct computation on all of the data in one big computation, which is impractically long or simply infeasible. D&R trades some accuracy for computational feasibility. The result is that almost any statistical or visualization method can be successfully applied to large complex data. This enables a deep, detailed analysis that does not risk losing important information in the data, which is feasible today only with small data.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Scalable Visualization and Model Building
  • 批准号:
    0937123
  • 项目类别:
    Standard Grant
  • 资助金额:
    $50.0万
  • 财政年份:
    2009
  • 负责人:
    William Cleveland
  • 依托单位:
Data Mining, Statistical Learning, and Data Visualization for Complex Data
  • 批准号:
    0532217
  • 项目类别:
    Standard Grant
  • 资助金额:
    $0.0万
  • 财政年份:
    2005
  • 负责人:
    William Cleveland
  • 依托单位:
国内基金
海外基金
Research on Quantum Field Theory without a Lagrangian Description
  • 批准号:
    24ZR1403900
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    SATOSHI NAWATA
  • 依托单位:
基于isomorph theory研究尘埃等离子体物理量的微观动力学机制
  • 批准号:
    12247163
  • 项目类别:
    专项项目
  • 资助金额:
    18.00万元
  • 批准年份:
    2022
  • 负责人:
    黄栋
  • 依托单位:
Toward a general theory of intermittent aeolian and fluvial nonsuspended sediment transport
  • 批准号:
    --
  • 项目类别:
    --
  • 资助金额:
    55万元
  • 批准年份:
    2022
  • 负责人:
    Thomas Pahtz
  • 依托单位:
英文专著《FRACTIONAL INTEGRALS AND DERIVATIVES: Theory and Applications》的翻译
  • 批准号:
    12126512
  • 项目类别:
    数学天元基金项目
  • 资助金额:
    12.0万元
  • 批准年份:
    2021
  • 负责人:
    李常品
  • 依托单位: