Larg-Scale Data Analytics: Methodologies and Applications
Larg-Scale Data Analytics: Methodologies and Applications
批准号:
RGPIN-2014-05721
负责人:
Ghodsi, Ali
金额:
$1.46万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2016
资助国家:
加拿大
项目状态:
已结题
起止时间:
2016-01-01 至 2017-12-31
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Recent years have witnessed the rise of the big data era in computing and storage systems. With the great advances in information and communication technology, hundreds of petabytes of data are generated, transferred, processed and stored every day. The availability of this overwhelming amount of structured and unstructured data creates an acute need to develop fast and accurate algorithms to discover useful information that is hidden in big data. One of the crucial problems in the big data era is the ability to represent the data and its underlying information in a succinct and interpretable format. Although different algorithms for clustering and dimensionality reduction can be used to summarize big data, these algorithms tend to learn representations whose meanings are difficult to interpret. For instance, the traditional clustering algorithms such as k-means tend to produce centroids which encode information about thousands of data instances, but the meanings of these centroids are hard to interpret. Even clustering methods that use data instances as prototypes, such as k-medoid, learn only one representative for each of the resulting clusters; these alone are not sufficient to capture the insights of the data instances in this cluster. In addition, using medoids as representatives implicitly assumes that the data points are distributed as clusters and that the number of those clusters is known ahead of time. This assumption is not true for all data sets. On the other hand, traditional dimensionality reduction algorithms such as Latent Semantic Analysis (LSA) tend to learn a few latent concepts in the feature space. Each of these concepts is represented by a dense vector which combines thousands of features with positive and negative weights. This makes it difficult for the data analyst to understand the meaning of these concepts. Even if the goal of representative selection is to learn a low-dimension embedding of data instances, learning dimensions whose meanings are easy to interpret allows the understanding of the results of data mining and machine learning algorithms, such as understanding the meanings of data clusters in the low-dimensional space. The acute need to summarize big data to a format that is informative for data analysts motivates the development of new algorithms to directly select a few representative data instances and/or features. This problem can be generally formulated as the selection of a subset of columns from a data matrix, which is formally known as the Column Subset Selection (CSS) problem. Although many algorithms have been proposed for tackling the CSS problem, most of these algorithms focus on randomly selecting a subset of columns with the goal of using these columns to obtain a low-rank approximation of the data matrix. In this case, these algorithms tend to select a relatively large number of columns. When the goal is to select a very few columns to be directly presented to a data analyst or indirectly used to interpret the results of other algorithms, the randomized CSS methods do not produce a meaningful subset of columns. On the other hand, deterministic algorithms for CSS, although more accurate, do not scale to work on big matrices with massively-distributed columns. We propose to address these limitations by developing a new framework that we call Data Downdate. A large variety of important problems in machine learning and data mining such as variable selection, mining representative patterns, and most notably sparse approximation are all special cases of this framework.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Beyond Deep Associative learning
-
批准号:RGPIN-2019-04824
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.68万
-
财政年份:2022
-
负责人:Ghodsi, Ali
-
依托单位:
Beyond Deep Associative learning
-
批准号:RGPIN-2019-04824
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.68万
-
财政年份:2021
-
负责人:Ghodsi, Ali
-
依托单位:
Beyond Deep Associative learning
-
批准号:RGPIN-2019-04824
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.68万
-
财政年份:2020
-
负责人:Ghodsi, Ali
-
依托单位:
Beyond Deep Associative learning
-
批准号:RGPIN-2019-04824
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.68万
-
财政年份:2019
-
负责人:Ghodsi, Ali
-
依托单位:
Larg-Scale Data Analytics: Methodologies and Applications
-
批准号:RGPIN-2014-05721
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.46万
-
财政年份:2018
-
负责人:Ghodsi, Ali
-
依托单位:
Larg-Scale Data Analytics: Methodologies and Applications
-
批准号:RGPIN-2014-05721
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.46万
-
财政年份:2017
-
负责人:Ghodsi, Ali
-
依托单位:
Larg-Scale Data Analytics: Methodologies and Applications
-
批准号:RGPIN-2014-05721
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.46万
-
财政年份:2015
-
负责人:Ghodsi, Ali
-
依托单位:
Larg-Scale Data Analytics: Methodologies and Applications
-
批准号:RGPIN-2014-05721
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.46万
-
财政年份:2014
-
负责人:Ghodsi, Ali
-
依托单位:
国内基金
海外基金
基于热量传递的传统固态发酵过程缩小(Scale-down)机理及调控
-
批准号:22108101
-
项目类别:青年科学基金项目(C类)
-
资助金额:30.0万元
-
批准年份:2021
-
负责人:靳光远
-
依托单位:
基于Multi-Scale模型的轴流血泵瞬变流及空化机理研究
-
批准号:31600794
-
项目类别:青年科学基金项目
-
资助金额:22.0万元
-
批准年份:2016
-
负责人:荆腾
-
依托单位:
针对Scale-Free网络的紧凑路由研究
-
批准号:60673168
-
项目类别:面上项目
-
资助金额:25.0万元
-
批准年份:2006
-
负责人:张国清
-
依托单位: