课题基金 / 基金详情

AF: DC: Collaborative Research: Pattern Matching for Massive Data Sets

AF: DC: Collaborative Research: Pattern Matching for Massive Data Sets
AF:DC:协作研究:海量数据集的模式匹配
批准号:
1017623
负责人:
Rahul Shah
金额:
$50.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2010
资助国家:
美国
项目状态:
已结题
起止时间:
2010-08-01 至 2014-07-31

项目摘要

项目成果

Rahul Shah的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
Pattern matching is a fundamental research field with applications in domains such as biological sequence alignment, web search engines and network intrusion detection. Given a pattern P and a text string T, the central problem is to find occurrences of P in T. When data becomes massive, we cannot assume that text can be stored in RAM. Pattern matching problems must now be considered with more appropriate models like external memory model, cache-oblivious model, streaming models, MapReduce paradigm and multi-core models. In many cases, a blend of models or newer, more appropriate models need to be developed keeping the practical aspects of the application in sight.The focus of this project is to develop efficient search algorithms and indexes when a data set resides on disks, on network storage, or is accessible only as an online stream. The data must be efficiently searchable even though it may be in compressed format. The project considers traditional pattern matching problem, as well as variants such as (i) approximate matching -- where the pattern may not exactly match a substring in T, (ii) online matching -- where the pattern(s) are known in advance and text comes as a stream, and (iii) string retrieval -- where instead of finding all the occurrences, the focus is on retrieving high ranking documents which contain one or more occurences of the query pattern. Issues of I/O efficiency and space utilization are central to this project. This involves developing suitable massive data set models, deriving optimal theoretical bounds and implementing practical tools. Methodologies include combinatorial and randomized methods in pattern matching, succinct data structures, top-k query processing and I/O efficient indexes.The project will build new, solid theoretical foundations in pattern matching, with direct applications to fields like databases and information retrieval. It will significantly drive forward current state of the art in web search engine technology (by impacting the way inverted indexes are used) and genome sequence alignment tools (e.g., BLAST). Tools and software developed during this project will be widely distributed to the research community. Some components will be incorporated into undergraduate and graduate algorithms course curricula as implementation projects.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
EAGER:CCF:AF:Sublinear Data Structures for Approximate Queries
  • 批准号:
    2137057
  • 项目类别:
    Standard Grant
  • 资助金额:
    $22.58万
  • 财政年份:
    2021
  • 负责人:
    Rahul Shah
  • 依托单位:
International Research Fellowship Program: Laser-Based Femtosecond X-Ray Development and Application
  • 批准号:
    0502281
  • 项目类别:
    Fellowship
  • 资助金额:
    $0.0万
  • 财政年份:
    2005
  • 负责人:
    Rahul Shah
  • 依托单位:
国内基金
海外基金
一种高变比DC/DC变换电路及控制方法的研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
    汪义旺
  • 依托单位:
单级式隔离型双向AC/DC变换器传导EMI建模及主动抑制方法研究
  • 批准号:
    2026JJ60451
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
    姜利
  • 依托单位:
复合功率处理的宽压高效双向DC-DC变换器及其模块化扩容方法研究
  • 批准号:
    2026JJ60452
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
    孙志峰
  • 依托单位:
疟原虫感染诱导DC表达ATG5抑制特异性CD4+Th1细胞活化的机制研究