课题基金 / 基金详情

III: Small: Probabilistic Hashing for Efficient Search Learning

III: Small: Probabilistic Hashing for Efficient Search Learning
III:小:用于高效搜索学习的概率哈希
批准号:
1360971
负责人:
Ping Li
金额:
$47.51万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2013
资助国家:
美国
项目状态:
已结题
起止时间:
2013-08-28 至 2018-08-31

项目摘要

项目成果

Ping Li的其他基金

相似基金

相关文献

中文摘要
翻译
许多应用涉及大量的高维数据集。例如,搜索行业通常处理数十亿个网页,其中每个页面通常表示为2^64维的二进制向量。在计算机视觉中,图像通常表示为数百万维的非二进制向量。能够有效地压缩、检索和挖掘这些数据集的算法具有很高的实际意义。将开发数学上严格和计算效率高的哈希方法,以大幅减少超高维数据集。这些算法将与各种学习技术,包括分类,聚类,近邻搜索,矩阵分解等项目的基础上建立和扩展minwise哈希,和b位minwise哈希这是标准的哈希技术在搜索应用程序。该项目旨在(i)严格分析b位minwise哈希,并开发,分析和应用显着更有效(ii)开发一个统一的概率散列框架,其基本上由一个排列和(最多)一个随机投影组成;(iii)在各种工程限制(储存空间、计算速度、索引能力、适应串流等)下,发展概括统计的统一理论。在此框架下开发的散列算法预计将比现有的流行算法,如随机投影和minwise散列更有效,更准确。 这个通用框架允许设计算法适应许多不同的数据类型(稀疏或密集数据,二进制或实值数据,静态或流数据),许多不同的工程需求(计算内积或lp距离,内核学习或线性学习),以及不同的存储要求。所提出的研究的预期结果包括用于处理超高维数据集的严格且计算高效的散列算法,将所得散列算法与用于分类、聚类、近邻搜索、奇异值分解、矩阵分解等的各种学习技术集成;以及对所得方法进行严格的实验评估(例如,TeraByte或潜在的PetaByte)数据,其数量级高达2^64维。更广泛的影响:从极高维数据中构建预测模型的有效方法可以影响许多依赖机器学习作为从数据中获取知识的主要方法的科学领域。PI的教育和外联工作旨在扩大妇女和代表性不足群体的参与。该项目产生的出版物、软件和数据集将免费传播给更广泛的科学界。
英文摘要
Numerous applications involve massive, high-dimensional datasets. For example, the search industry routinely deals with billions of web pages, where each page is often represented as a binary vector in 2^64 dimensions. In computer vision, images are often represented as non-binary vectors in millions of dimensions. Algorithms which are capable of efficiently compressing, retrieving, and mining these datasets are of high practical importance. Mathematically rigorous and computationally efficient hashing methods will be developed to dramatically reduce ultra-high-dimensional datasets. These algorithms will be integrated with a variety of learning techniques including classification, clustering, near-neighbor search, matrix factorizations, etc. The project builds on and extends minwise hashing, and b-bit minwise hashing which are standard hashing techniques in search applications. The project aims to (i) rigorously analyze b-bit minwise hashing and develop, analyze, and apply significantly more efficient (and more accurate) to problems in search and learning; (ii) develop a unified framework of probabilistic hashing which essentially consists of one permutation followed by (at most) one random projection; (iii) develop a unified theory of summary statistics under a variety of engineering constraints (storage space, computational speed, indexing capability, adaptation to streaming, etc.). Hashing algorithms developed under this framework are expected to be substantially much more efficient and more accurate than existing popular algorithms such as random projections and minwise hashing. This general framework allows the design algorithms to accommodate many different data types (sparse or dense data, binary or real-valued data, static or streaming data), many different engineering needs (computing inner products or lp distances, kernel learning or linear learning), and different storage requirements. Anticipated results of the proposed research include rigorous and computationally efficient hashing algorithms for dealing with ultra-high-dimensional datasets, the integration of the resulting hashing algorithms into with a variety of learning techniques for classification, clustering, near-neighbor search, singular value decompositions, matrix factorization, etc; and rigorous experimental evaluation of the resulting methods on big (e.g., TeraByte or potentially PetaByte) data of the order of up to 2^64 dimensions. Broader Impacts: Effective approaches to building predictive models from extremely high dimensional data can impact many areas of science that rely on machine learning as the primary methodology for knowledge acquisition from data. The PI's education and outreach efforts aim to broaden the participation of women and underrepresented groups. The publications, software, and datasets resulting from the project will be freely disseminated to the larger scientific community.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: Study of A- and B-class dye-decolorizing peroxidases (DyPs): From molecular mechanisms to applications in dye removal and lignin degradation
  • 批准号:
    1807532
  • 项目类别:
    Standard Grant
  • 资助金额:
    $46.18万
  • 财政年份:
    2018
  • 负责人:
    Ping Li
  • 依托单位:
Efficient Data Reduction and Summarization
  • 批准号:
    1444124
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $10.49万
  • 财政年份:
    2014
  • 负责人:
    Ping Li
  • 依托单位:
Neurocognitive Mechanisms of Second Language Learning: Role of Learning Context and Cognitive Functions
BIGDATA: Small: DA: A Random Projection Approach
  • 批准号:
    1419210
  • 项目类别:
    Standard Grant
  • 资助金额:
    $44.74万
  • 财政年份:
    2013
  • 负责人:
    Ping Li
  • 依托单位:
国内基金
海外基金
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
  • 依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2022
  • 负责人:
    张祥忠
  • 依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
  • 批准号:
    31972324
  • 项目类别:
    面上项目
  • 资助金额:
    58.0万元
  • 批准年份:
    2019
  • 负责人:
    高学文
  • 依托单位: