CAREER: Scalable Algorithms for Large-Scale Data Mining
CAREER: Scalable Algorithms for Large-Scale Data Mining
批准号:
0093404
负责人:
Inderjit Dhillon
金额:
$0.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2001
资助国家:
美国
项目状态:
已结题
起止时间:
2001-06-01 至 2008-05-31
中文摘要
数字数据可以以多种形式出现;它可以作为带有数字字段的数据库记录、原始文本文档或图像文件或网站流量日志文件出现。数据挖掘是自动发现数据中有趣的模式、关联、变化、异常、规则以及统计上重要的结构和事件。数据的一个关键特征,通常是压倒性的特征,是其绝对的数量。快速扩展的互联网已经包含超过10亿个网页,典型的仓库和Web流量数据可以占用TB的磁盘空间。很明显,数据挖掘工具必须是有效的和可扩展的,如果他们要服务于任何实际目的。并行计算可以帮助满足这些大型数据集对计算周期和内存存储的需求。本项目的主要重点是开发可扩展的解决方案,用于大规模数据分析。其主要目标是探索和开发高效、并行的数学和统计方法,以挖掘大型数据集并及时提供结果。特别是,新的聚类技术,分区数据到不相交的分区,概念分解的新方法降维,改进计算的主成分分析,有效的分类方案折叠在新到达的未标记的数据到已知的类,和有效的可视化多维数据将进行调查。另一个重点是使开发的数据分析工具适应文本挖掘的应用领域,构建一个完全并行的文本挖掘系统,该系统能够(a)有效地对文本数据进行数字化预处理,(B)对大型未标记文档集合进行聚类,(c)将未标记文档分类到已知的概念层次中,(d)可视化文档词之间的关系。该系统将使用户能够方便地浏览、吸收、搜索和组织大量文件的内容;我们希望在一个128个处理器的工作站集群上处理多达1亿份文件。我们开发的许多文本挖掘算法将随着数据的大小线性扩展。在这种情况下,避免I/O瓶颈,利用现代处理器的内存层次结构和隐藏网络延迟变得非常重要。教育计划包括三个组成部分:(i)强调本科和研究生教育中的科学方法的教学理念,通过结合新技术进行课堂教学和基于网络的离线教学; ㈡注重多学科教育,致力于开发集中的面向网络的初级读本,以便使学生迅速熟悉所需的预科课程; ㈢编制两门课程的课程;第一个是为非CS本科生开设的科学计算课程,作为UT Austin新的“计算元素”计划的一部分,第二个是为研究生开设的大规模数据挖掘新课程。
英文摘要
Digital data can occur in diverse forms; it may occur as database records with numerical fields, as raw text documents or image files, or as website traffic log files. Data mining is the automatic discovery of interesting patterns, associations,changes, anomalies, rules, and statistically significant structures and events in data. A key feature, often an overwhelming feature, of the data is its sheer magnitude. The rapidly expanding internet already contains more than 1 billion web pages, and typical warehouse and web traffic data can occupy terabytes of disk space. It is clear that data mining tools must be efficient and scalable if they are to serve any practical purpose. Parallel computing can help in satisfying the demands on computing cycles and memory storage imposed by these large data sets.The main focus of this project is to develop scalable solutions for large-scale data analysis. The main thrust isin exploring and developing efficient, parallel, mathematical and statistical methods that can mine large data sets and deliver results in a timely manner. In particular, new clustering techniques that partition data intodisjoint partitions, the new method of concept decompositions for dimensionality reduction, improved computation of principal components analysis, efficient classification schemes for folding in newly arriving unlabeled data into known classes, and effective visualization of multidimensional data will be investigated. Another focus is to adapt the data analyses tools developed to the application area of text mining.A completely parallel text mining system that is capable of (a) efficient preprocessing of text data intonumerical data, (b) clustering large unlabeled document collections, (c) classifying unlabeled documents into a known concept hierarchy and (d) visualization of document & word relationships will be built. This system will allow the user to easily navigate, assimilate, search and organize the contents of very large document collections; we hope to process up to 100 million documents on a 128-processor cluster of workstations. Many of the text mining algorithms we develop will scale linearly with the size of the data. In this scenario, it becomes important to avoid I/O bottlenecks, exploit memory hierarchies of modern processors and hide network latencies.The educational plan consists of three components: (i) a teaching philosophy that emphasizesthe scientific method in undergraduate and graduate education, by incorporating new technologies for in-class and web-based offline instruction; (ii) a focus on multidisciplinary education with a commitment to develop centralized web-oriented primers designed to quickly acquaint students with desired pre-requisites; and (iii) curriculum development for two courses; the first, a scientific computing course for non-CS undergraduates as part of UT Austin's new "Elements of Computing" program, and the second, a new course on large-scale data mining for graduate students.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
BIGDATA: Collaborative Research: F: Nomadic Algorithms for Machine Learning in the Cloud
-
批准号:1546452
-
项目类别:Standard Grant
-
资助金额:$61.04万
-
财政年份:2016
-
负责人:Inderjit Dhillon
-
依托单位:
I-Corps: Faster than Light Big Data Analytics
-
批准号:1507631
-
项目类别:Standard Grant
-
资助金额:$5.0万
-
财政年份:2015
-
负责人:Inderjit Dhillon
-
依托单位:
AF:Small: Divide-and-Conquer Numerical Methods for Analysis of Massive Data Sets
-
批准号:1320746
-
项目类别:Standard Grant
-
资助金额:$49.1万
-
财政年份:2013
-
负责人:Inderjit Dhillon
-
依托单位:
AF: Small: Fast and Memory-Efficient Dimensionality Reduction for Massive Networks
-
批准号:1117055
-
项目类别:Standard Grant
-
资助金额:$36.0万
-
财政年份:2011
-
负责人:Inderjit Dhillon
-
依托单位:
Non-Negative Matrix and Tensor Approximations: Algorithms, Software and Applications
-
批准号:0728879
-
项目类别:Standard Grant
-
资助金额:$25.0万
-
财政年份:2007
-
负责人:Inderjit Dhillon
-
依托单位:
Novel Matrix Problems in Modern Applications
-
批准号:0431257
-
项目类别:Standard Grant
-
资助金额:$0.0万
-
财政年份:2004
-
负责人:Inderjit Dhillon
-
依托单位:
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
-
批准号:--
-
项目类别:合作创新研究团队
-
资助金额:--
-
批准年份:2024
-
负责人:姚韬
-
依托单位: