课题基金 / 基金详情

Large-Scale Matrix Computation Problems in Information Retrieval and Datamining

Large-Scale Matrix Computation Problems in Information Retrieval and Datamining
信息检索和数据挖掘中的大规模矩阵计算问题
批准号:
9901986
负责人:
Hongyuan Zha
金额:
$22.93万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
1999
资助国家:
美国
项目状态:
已结题
起止时间:
1999-08-15 至 2003-07-31

项目摘要

项目成果

Hongyuan Zha的其他基金

相似基金

相关文献

中文摘要
翻译
本研究的重点是开发,分析和实现的快速和内存效率的算法,解决了几个矩阵计算问题所产生的信息检索和数据挖掘。 大型和/或稀疏矩阵的奇异值分解(SVD)在这项研究中起着至关重要的作用:它被用来,一方面,降维,以实现高算法效率,另一方面,降噪,以提高检索精度。 重点是探索这些应用程序所产生的矩阵的内在结构,以开发其SVD计算的快速算法。具体而言,本研究的目标是:1)为信息检索、超文本链接分析和数据挖掘等领域的潜在语义索引(LSI)和相关的谱分析方法奠定理论基础; 2)开发和实现用于部分SVD和稀疏低秩近似计算的高性能算法,该算法可以扩展到非常大的文本语料库和数据库。的统计模型已经被开发来表示术语和潜在概念之间的关系。 这个模型导致了一个“低秩加移位”的结构,它近似地满足于术语-文档矩阵的叉积。这种结构产生了一个更准确的更新计划LSI,它也导致了一个高度并行的分治法计算大型稀疏术语文档矩阵的部分SVD。本项目将1)进一步发展我们的基于子空间的模型,并深入了解部分奇异值分解及其稀疏变化在文本检索、链接分析和数据挖掘应用中的有效性; 2)进一步探索术语-文档矩阵和链接矩阵的低秩加移位结构,以开发快速和节省内存的数值算法,用于在线性时间内计算其部分奇异值分解; 3)探索其他有效的算法来寻找用于Web链接分析和数据挖掘应用的稠密二部子图; 4)在WWW和大型商业数据库中生成的文本语料库上对算法进行了测试,WWW的数据挖掘应用和链接分析为研究大规模计算的算法问题和并行实现问题提供了一个完美的平台。 本研究试图整合基于统计和矩阵扰动理论的理论研究,使用计算线性代数方法的算法开发,以及对现实世界的文本语料库和从WWW搜索引擎和大型商业数据库获得的大型数据集的实验。
英文摘要
This research focuses on the development, analysis and implementation of fast and memory-efficient algorithms for solving several matrix computation problems arising from information retrieval and data mining. The singular value decomposition (SVD) of large and/or sparse matrices plays an essential role in this research: it is used, on the one hand, for dimension reduction to achieve high algorithmic efficiency, and on the other for noise reduction to improve retrieval accuracy. The emphasis is on exploring the intrinsic structures of the matrices arising from those applications to develop fast algorithms for their SVD computation. Specifically, the goals of this research are 1) the development of theoretical foundation for latent semantic indexing (LSI) and related spectral analytical methods for information retrieval, hypertext link analysis and data mining; 2) the development and implementation of high performance algorithms for partial SVD and sparse low-rank approximation computation that can scale to very large text corpora and databases.A subspace-based statistical model has been previously developed to represent the relations between terms and latent concepts. This model has led to a "low-rank-plus-shift" structure that is approximately satisfied by the cross-product of the term-document matrices. This structure gives rise to a more accurate updating scheme for LSI and it also leads to a highly parallel divide-and-conquer method for computing the partial SVD of large sparse term-document matrices. This project will 1) further develop our subspace-based model and gain deeper understanding of the effectiveness of partial SVD and their sparse variations in text retrieval, link analysis and data mining applications; 2) further explore the low-rank-plus-shift structures of the term-document matrices and link matrices to develop fast and memory-efficient numerical algorithms for the computation of their partial SVD in linear time; 3) explore other efficient heuristics for finding dense bipartite subgraphs used in Web link analysis and data mining applications; and 4) test the algorithms on text corpora generated from WWW and large commercial databases.No computation of SVD on the order of several million has been attempted before, and LSI, data mining applications and link analysis for WWW provides a perfect platform to investigate algorithmic issues and parallel implementation issues for large-scale computation. This research attempts to integrate theoretical investigation based on statistics and matrix perturbation theory, algorithmic development using computational linear algebra methodologies, and experimentation on real-world text corpora and large datasets obtained from WWW search engines and large commercial databases.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: CDS&E-MSS: Robust Algorithms for Interpolation and Extrapolation in Manifold Learning
  • 批准号:
    1317372
  • 项目类别:
    Standard Grant
  • 资助金额:
    $17.0万
  • 财政年份:
    2013
  • 负责人:
    Hongyuan Zha
  • 依托单位:
III: Small: Exploring Social and Behavioral Contexts for Information Retrieval
  • 批准号:
    1116886
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $49.6万
  • 财政年份:
    2011
  • 负责人:
    Hongyuan Zha
  • 依托单位:
III: EAGER: Learning Evaluation Metrics for Information Retrieval
  • 批准号:
    1049694
  • 项目类别:
    Standard Grant
  • 资助金额:
    $20.0万
  • 财政年份:
    2010
  • 负责人:
    Hongyuan Zha
  • 依托单位:
Computational Methods for Nonlinear Dimension Reduction
  • 批准号:
    0736328
  • 项目类别:
    Standard Grant
  • 资助金额:
    $0.0万
  • 财政年份:
    2007
  • 负责人:
    Hongyuan Zha
  • 依托单位:
国内基金
海外基金
基于热量传递的传统固态发酵过程缩小(Scale-down)机理及调控
  • 批准号:
    22108101
  • 项目类别:
    青年科学基金项目(C类)
  • 资助金额:
    30.0万元
  • 批准年份:
    2021
  • 负责人:
    靳光远
  • 依托单位:
基于Multi-Scale模型的轴流血泵瞬变流及空化机理研究
  • 批准号:
    31600794
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    22.0万元
  • 批准年份:
    2016
  • 负责人:
    荆腾
  • 依托单位:
针对Scale-Free网络的紧凑路由研究