课题基金 / 基金详情

Larg-Scale Data Analytics: Methodologies and Applications

Larg-Scale Data Analytics: Methodologies and Applications
大规模数据分析:方法和应用
批准号:
RGPIN-2014-05721
负责人:
Ghodsi, Ali
金额:
$1.46万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2014
资助国家:
加拿大
项目状态:
已结题
起止时间:
2014-01-01 至 2015-12-31

项目摘要

项目成果

Ghodsi, Ali的其他基金

相似基金

相关文献

中文摘要
翻译
近年来,计算和存储系统领域的大数据时代正在兴起。随着信息和通信技术的巨大进步,每天都会产生、传输、处理和存储数以百计的数据。海量的结构化和非结构化数据的可获得性迫切需要开发快速、准确的算法来发现隐藏在大数据中的有用信息。大数据时代的关键问题之一是以简洁和可解释的格式表示数据及其底层信息的能力。尽管可以使用不同的聚类和降维算法来总结大数据,但这些算法倾向于学习含义难以理解的表示法。例如,传统的聚类算法,如k-Means,往往会产生编码数千个数据实例信息的质心,但这些质心的含义很难解释。即使是使用数据实例作为原型的集群方法,如k-medoid,也只为每个结果集群学习一个代表;这些本身不足以捕获该集群中数据实例的洞察力。此外,使用medoid作为代表,隐含地假设数据点是以簇的形式分布的,并且这些簇的数量是提前知道的。这一假设并不适用于所有数据集。另一方面,潜在语义分析(LSA)等传统降维算法倾向于学习特征空间中的一些潜在概念。这些概念中的每一个都由一个密集的向量表示,该向量结合了数千个具有正负权重的特征。这使得数据分析师很难理解这些概念的含义。即使代表性选择的目标是学习数据实例的低维嵌入,其含义易于解释的学习维度也允许理解数据挖掘和机器学习算法的结果,例如理解低维空间中数据簇的含义。将大数据汇总为一种对数据分析师有用的格式的迫切需求促使新算法的开发,以直接选择几个具有代表性的数据实例和/或特征。该问题通常可以表示为从数据矩阵中选择列的子集,其正式名称为列子集选择(CSS)问题。虽然已经提出了许多算法来解决CSS问题,但这些算法中的大多数集中于随机选择列的子集,目的是使用这些列来获得数据矩阵的低级近似。在这种情况下,这些算法倾向于选择相对较大数量的列。当目标是选择极少数列直接呈现给数据分析师或间接用于解释其他算法的结果时,随机化的CSS方法不会产生有意义的列子集。另一方面,用于CSS的确定性算法虽然更精确,但不能扩展以处理具有大量分布的列的大矩阵。我们建议通过开发一个称为Data Downdate的新框架来解决这些限制。机器学习和数据挖掘中的许多重要问题,如变量选择、挖掘代表性模式,以及最值得注意的稀疏近似,都是该框架的特例。
英文摘要
Recent years have witnessed the rise of the big data era in computing and storage systems. With the great advances in information and communication technology, hundreds of petabytes of data are generated, transferred, processed and stored every day. The availability of this overwhelming amount of structured and unstructured data creates an acute need to develop fast and accurate algorithms to discover useful information that is hidden in big data. One of the crucial problems in the big data era is the ability to represent the data and its underlying information in a succinct and interpretable format. Although different algorithms for clustering and dimensionality reduction can be used to summarize big data, these algorithms tend to learn representations whose meanings are difficult to interpret. For instance, the traditional clustering algorithms such as k-means tend to produce centroids which encode information about thousands of data instances, but the meanings of these centroids are hard to interpret. Even clustering methods that use data instances as prototypes, such as k-medoid, learn only one representative for each of the resulting clusters; these alone are not sufficient to capture the insights of the data instances in this cluster. In addition, using medoids as representatives implicitly assumes that the data points are distributed as clusters and that the number of those clusters is known ahead of time. This assumption is not true for all data sets. On the other hand, traditional dimensionality reduction algorithms such as Latent Semantic Analysis (LSA) tend to learn a few latent concepts in the feature space. Each of these concepts is represented by a dense vector which combines thousands of features with positive and negative weights. This makes it difficult for the data analyst to understand the meaning of these concepts. Even if the goal of representative selection is to learn a low-dimension embedding of data instances, learning dimensions whose meanings are easy to interpret allows the understanding of the results of data mining and machine learning algorithms, such as understanding the meanings of data clusters in the low-dimensional space. The acute need to summarize big data to a format that is informative for data analysts motivates the development of new algorithms to directly select a few representative data instances and/or features. This problem can be generally formulated as the selection of a subset of columns from a data matrix, which is formally known as the Column Subset Selection (CSS) problem. Although many algorithms have been proposed for tackling the CSS problem, most of these algorithms focus on randomly selecting a subset of columns with the goal of using these columns to obtain a low-rank approximation of the data matrix. In this case, these algorithms tend to select a relatively large number of columns. When the goal is to select a very few columns to be directly presented to a data analyst or indirectly used to interpret the results of other algorithms, the randomized CSS methods do not produce a meaningful subset of columns. On the other hand, deterministic algorithms for CSS, although more accurate, do not scale to work on big matrices with massively-distributed columns. We propose to address these limitations by developing a new framework that we call Data Downdate. A large variety of important problems in machine learning and data mining such as variable selection, mining representative patterns, and most notably sparse approximation are all special cases of this framework.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Beyond Deep Associative learning
  • 批准号:
    RGPIN-2019-04824
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.68万
  • 财政年份:
    2022
  • 负责人:
    Ghodsi, Ali
  • 依托单位:
Beyond Deep Associative learning
  • 批准号:
    RGPIN-2019-04824
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.68万
  • 财政年份:
    2021
  • 负责人:
    Ghodsi, Ali
  • 依托单位:
Beyond Deep Associative learning
  • 批准号:
    RGPIN-2019-04824
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.68万
  • 财政年份:
    2020
  • 负责人:
    Ghodsi, Ali
  • 依托单位:
Beyond Deep Associative learning
  • 批准号:
    RGPIN-2019-04824
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.68万
  • 财政年份:
    2019
  • 负责人:
    Ghodsi, Ali
  • 依托单位:
国内基金
海外基金
基于热量传递的传统固态发酵过程缩小(Scale-down)机理及调控
  • 批准号:
    22108101
  • 项目类别:
    青年科学基金项目(C类)
  • 资助金额:
    30.0万元
  • 批准年份:
    2021
  • 负责人:
    靳光远
  • 依托单位:
基于Multi-Scale模型的轴流血泵瞬变流及空化机理研究
  • 批准号:
    31600794
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    22.0万元
  • 批准年份:
    2016
  • 负责人:
    荆腾
  • 依托单位:
针对Scale-Free网络的紧凑路由研究