Larg-Scale Data Analytics: Methodologies and Applications
Larg-Scale Data Analytics: Methodologies and Applications
批准号:
RGPIN-2014-05721
负责人:
Ghodsi, Ali
金额:
$1.46万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2018
资助国家:
加拿大
项目状态:
已结题
起止时间:
2018-01-01 至 2019-12-31
中文摘要
近年来,大数据时代在计算和存储系统中兴起。随着信息和通信技术的巨大进步,每天都会产生、传输、处理和存储数百pb的数据。由于大量的结构化和非结构化数据的存在,迫切需要开发快速准确的算法来发现隐藏在大数据中的有用信息。大数据时代的关键问题之一是能够以简洁和可解释的格式表示数据及其底层信息。虽然可以使用不同的聚类和降维算法来总结大数据,但这些算法往往学习难以解释含义的表示。例如,传统的聚类算法(如k-means)倾向于产生包含数千个数据实例信息的质心,但这些质心的含义很难解释。即使是使用数据实例作为原型的聚类方法,比如k-medoid,也只能为每个聚类学习一个代表;仅凭这些还不足以捕获集群中数据实例的洞察力。此外,使用medioid作为代表隐含地假设数据点是作为集群分布的,并且这些集群的数量是提前已知的。这个假设并不适用于所有数据集。另一方面,传统的降维算法如Latent Semantic Analysis (LSA)倾向于学习特征空间中的一些潜在概念。这些概念中的每一个都由一个密集向量表示,该向量结合了数千个具有正权和负权的特征。这使得数据分析人员很难理解这些概念的含义。即使代表性选择的目标是学习数据实例的低维嵌入,学习含义易于解释的维度也可以理解数据挖掘和机器学习算法的结果,例如理解低维空间中数据簇的含义。迫切需要将大数据总结为一种对数据分析人员有用的格式,这激发了新算法的开发,以直接选择一些具有代表性的数据实例和/或特征。这个问题通常可以表示为从数据矩阵中选择列的子集,其正式名称为列子集选择(CSS)问题。尽管已经提出了许多算法来解决CSS问题,但这些算法中的大多数都集中在随机选择列的子集,目的是使用这些列来获得数据矩阵的低秩近似值。在这种情况下,这些算法倾向于选择相对较多的列。当目标是选择很少的列直接呈现给数据分析师或间接用于解释其他算法的结果时,随机化CSS方法不会产生有意义的列子集。另一方面,CSS的确定性算法虽然更精确,但不能扩展到具有大量分布列的大矩阵。我们建议通过开发一个我们称之为Data Downdate的新框架来解决这些限制。机器学习和数据挖掘中的大量重要问题,如变量选择、挖掘代表性模式以及最显著的稀疏近似都是该框架的特殊情况。
英文摘要
Recent years have witnessed the rise of the big data era in computing and storage systems. With the great advances in information and communication technology, hundreds of petabytes of data are generated, transferred, processed and stored every day. The availability of this overwhelming amount of structured and unstructured data creates an acute need to develop fast and accurate algorithms to discover useful information that is hidden in big data. One of the crucial problems in the big data era is the ability to represent the data and its underlying information in a succinct and interpretable format. Although different algorithms for clustering and dimensionality reduction can be used to summarize big data, these algorithms tend to learn representations whose meanings are difficult to interpret. For instance, the traditional clustering algorithms such as k-means tend to produce centroids which encode information about thousands of data instances, but the meanings of these centroids are hard to interpret. Even clustering methods that use data instances as prototypes, such as k-medoid, learn only one representative for each of the resulting clusters; these alone are not sufficient to capture the insights of the data instances in this cluster. In addition, using medoids as representatives implicitly assumes that the data points are distributed as clusters and that the number of those clusters is known ahead of time. This assumption is not true for all data sets. On the other hand, traditional dimensionality reduction algorithms such as Latent Semantic Analysis (LSA) tend to learn a few latent concepts in the feature space. Each of these concepts is represented by a dense vector which combines thousands of features with positive and negative weights. This makes it difficult for the data analyst to understand the meaning of these concepts. Even if the goal of representative selection is to learn a low-dimension embedding of data instances, learning dimensions whose meanings are easy to interpret allows the understanding of the results of data mining and machine learning algorithms, such as understanding the meanings of data clusters in the low-dimensional space. The acute need to summarize big data to a format that is informative for data analysts motivates the development of new algorithms to directly select a few representative data instances and/or features. This problem can be generally formulated as the selection of a subset of columns from a data matrix, which is formally known as the Column Subset Selection (CSS) problem. Although many algorithms have been proposed for tackling the CSS problem, most of these algorithms focus on randomly selecting a subset of columns with the goal of using these columns to obtain a low-rank approximation of the data matrix. In this case, these algorithms tend to select a relatively large number of columns. When the goal is to select a very few columns to be directly presented to a data analyst or indirectly used to interpret the results of other algorithms, the randomized CSS methods do not produce a meaningful subset of columns. On the other hand, deterministic algorithms for CSS, although more accurate, do not scale to work on big matrices with massively-distributed columns. We propose to address these limitations by developing a new framework that we call Data Downdate. A large variety of important problems in machine learning and data mining such as variable selection, mining representative patterns, and most notably sparse approximation are all special cases of this framework.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Beyond Deep Associative learning
-
批准号:RGPIN-2019-04824
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.68万
-
财政年份:2022
-
负责人:Ghodsi, Ali
-
依托单位:
Beyond Deep Associative learning
-
批准号:RGPIN-2019-04824
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.68万
-
财政年份:2021
-
负责人:Ghodsi, Ali
-
依托单位:
Beyond Deep Associative learning
-
批准号:RGPIN-2019-04824
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.68万
-
财政年份:2020
-
负责人:Ghodsi, Ali
-
依托单位:
Beyond Deep Associative learning
-
批准号:RGPIN-2019-04824
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.68万
-
财政年份:2019
-
负责人:Ghodsi, Ali
-
依托单位:
Larg-Scale Data Analytics: Methodologies and Applications
-
批准号:RGPIN-2014-05721
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.46万
-
财政年份:2017
-
负责人:Ghodsi, Ali
-
依托单位:
Larg-Scale Data Analytics: Methodologies and Applications
-
批准号:RGPIN-2014-05721
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.46万
-
财政年份:2016
-
负责人:Ghodsi, Ali
-
依托单位:
Larg-Scale Data Analytics: Methodologies and Applications
-
批准号:RGPIN-2014-05721
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.46万
-
财政年份:2015
-
负责人:Ghodsi, Ali
-
依托单位:
Larg-Scale Data Analytics: Methodologies and Applications
-
批准号:RGPIN-2014-05721
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.46万
-
财政年份:2014
-
负责人:Ghodsi, Ali
-
依托单位:
国内基金
海外基金
基于热量传递的传统固态发酵过程缩小(Scale-down)机理及调控
-
批准号:22108101
-
项目类别:青年科学基金项目(C类)
-
资助金额:30.0万元
-
批准年份:2021
-
负责人:靳光远
-
依托单位:
基于Multi-Scale模型的轴流血泵瞬变流及空化机理研究
-
批准号:31600794
-
项目类别:青年科学基金项目
-
资助金额:22.0万元
-
批准年份:2016
-
负责人:荆腾
-
依托单位:
针对Scale-Free网络的紧凑路由研究
-
批准号:60673168
-
项目类别:面上项目
-
资助金额:25.0万元
-
批准年份:2006
-
负责人:张国清
-
依托单位: