课题基金 / 基金详情

Statistical inference for distributed datasets

Statistical inference for distributed datasets
分布式数据集的统计推断
批准号:
RGPIN-2016-06296
负责人:
Plante, JeanFrançois
金额:
$1.09万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2016
资助国家:
加拿大
项目状态:
已结题
起止时间:
2016-01-01 至 2017-12-31

项目摘要

项目成果

Plante, JeanFrançois的其他基金

相似基金

相关文献

中文摘要
翻译
技术使我们能够存储和描述大量数据。这些数据通常是关于人类的,记录个人生活,让其他人比自己更了解客户,并提供对社会动态的特殊洞察。大数据时代已经到来,如此大量的信息不能存储在个人电脑上,而是需要相互连接的机器集群。这种大规模的设置通常基于分布式文件系统(如Hadoop),因为单个驱动器不可能存储那么多信息。使用“分布式数据”,单个处理器无法访问整个数据集,相反,大量处理器每个只能访问数据的一部分。
英文摘要
Technology allows us to store and describe massive amounts of data. These data are often about humans, keeping a personal trace of individual lives, allowing others to know a customer better than himself, and providing a special insight into social dynamics. The era of big data is here, and such massive amounts of information cannot be stored on a personal computer, but rather require clusters of machines that are interconnected. Such large scale setups are typically based on distributed file system (such as Hadoop) because no single drive could possibly store that much information. With “distributed data”, a single processor is unable to access the whole dataset, but instead, a large number of processors each have access to only a part of the data. Statistical analyses can extract knowledge from big data, but most statistical models were designed for smaller datasets, assuming that the whole data was available from one computer. This assumption does not hold for large-scale samples and only a few statistical methods are straightforward to adapt. For the vast majority of statistical tools, unless one is willing to analyse only a randomly selected fraction of a massive dataset, new innovative solutions are required. The main objective of this research program is to adapt statistical tools to the reality of distributed data. Many statistical procedures are based on the estimated law of a random variable, but in a distributed environment, the communication between the different computing nodes is a scarce resource. Sharing all of the data is not an option and a compromise must be struck between precision and communication costs. Building an estimate for the law of a variable of interest is a fundamental challenge and this research program proposes a number of strategies to do so within the constraints of a distributed architecture. Once the properties of the proposed estimates are well established, they will be used to generalize statistical procedures as diverse as maximum likelihood estimation, goodness-of-fit tests and resampling procedures including the bootstrap. Moving things one step further, strategies are also proposed to infer the multivariate dependence structure of the data, the infamous copula. By developing statistical methodology for distributed (big) data, this research project will provide tools that are highly needed in all sciences just as much as in industrial and business applications who all share the need to extract knowledge from data.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Statistical inference for distributed datasets
  • 批准号:
    RGPIN-2016-06296
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.09万
  • 财政年份:
    2021
  • 负责人:
    Plante, JeanFrançois
  • 依托单位:
Statistical inference for distributed datasets
  • 批准号:
    RGPIN-2016-06296
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.09万
  • 财政年份:
    2020
  • 负责人:
    Plante, JeanFrançois
  • 依托单位:
Statistical inference for distributed datasets
  • 批准号:
    RGPIN-2016-06296
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.09万
  • 财政年份:
    2019
  • 负责人:
    Plante, JeanFrançois
  • 依托单位:
Statistical inference for distributed datasets
  • 批准号:
    RGPIN-2016-06296
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.09万
  • 财政年份:
    2018
  • 负责人:
    Plante, JeanFrançois
  • 依托单位:
海外基金