课题基金 / 基金详情

Development of Efficient Data Mining Systems for Large Semi-Structured Text Data

Development of Efficient Data Mining Systems for Large Semi-Structured Text Data
大型半结构化文本数据的高效数据挖掘系统开发
批准号:
11558040
负责人:
ARIMURA Hiroki
金额:
$6.27万
依托单位:
依托单位国家:
日本
项目类别:
Grant-in-Aid for Scientific Research (B)
财政年份:
1999
资助国家:
日本
项目状态:
已结题
起止时间:
1999 至 2001

项目摘要

项目成果

ARIMURA Hiroki的其他基金

相似基金

相关文献

中文摘要
翻译
这个研究项目的目标是设计一个高效的半自动工具,支持人类从大型非结构化和半结构化的文本数据中发现。为了实现这一目标,我们从以下三个方向进行了研究.文本挖掘的核心过程是模式发现。我们采用了优化模式发现的框架,并开发了高效和强大的文本挖掘算法,发现简单的组合模式,从大型非结构化文本。为了实现这些算法,我们开发了一个适合于文本挖掘的基于后缀数组的文本索引结构。基于这些技术,我们实现了一个原型系统,并对Web数据进行了计算机实验.文本的另一个重要技术是高效的模式匹配。作为一个理论框架,我们提出了一个统一的框架,称为拼贴系统,实现各种基于字典的压缩方法。我们开发了Knuth-Morris-Pratt型和Byer-Moore型模式匹配算法,采用这个框架。并将该框架应用于Byte-Pair-Encoding压缩方法和Sequitur中,前者产生了最快的压缩模式匹配算法.文本挖掘的最后一个过程是信息抽取。从理论的角度,我们首先形式化的信息提取问题,从半结构化数据,然后给这样的任务的能力和局限性的理论分析。然后,我们开发了有效的信息提取算法的各种类型的提取规则,包括树包装和对冲模式,并评估他们通过实验在现实生活中的半结构化数据在互联网上。
英文摘要
The goal of this research project is to devise an efficient semi-automatic tool that supports human discovery from large unstructured and semi-structured text data. To achieve this goal, we studied in the following three directions.1. The central process of text mining is pattern discovery. We employed the framework of optimized pattern discovery, and developed effcient and robust text mining algorithms that find simple combinatorial patterns from large unstructured texts. To implement these algorithms, we developed a text index structure based on the suffix arrays suitable for text mining. Based on these technologies, we implemented a prototype system and run computer experiments on Web data.2. Another important technology for text is efficient pattern matching. As a theoretical framework, we proposed a unified framework, called Collage system, for realizing various dictionary-based compression methods. We developed both Knuth-Morris-Pratt type and Byer-Moore type pattern matching algorithms employing this framework. We also applied this framework to Byte-Pair-Encoding compression method and Sequitur, the former of which yields the fastest compressed pattern matching algorithm.3. Final process of text mining is information extraction. From theoretical point of view, we first formalize the information extraction problem from semi-structured data, and then gave theoretical analysis of the power and the limitation of such tasks. Then, we developed efficient information extraction algorithms for various types of extraction rules including tree wrappers and hedge patterns and evaluate them through experiments on real-life semi-structured data on the internet.
期刊论文(104)
专著(0)
科研奖励(0)
会议论文
DOI: --
发表时间:
期刊:
影响因子: --
作者: []
通讯作者:
DOI: --
发表时间:
期刊:
影响因子: --
作者: []
通讯作者:
M.Takeda et al.: "Mining from Literary Texts : Pattern Discovery and Similarity Computation"Lecture Notes in Computer Science. 2281. 520-533 (2002)
M.Takeda 等人:“从文学文本中挖掘:模式发现和相似性计算”计算机科学讲义。
DOI: --
发表时间:
期刊:
影响因子: --
作者: []
通讯作者:
H.Arimura et al.: "Efficient Learning of Semi-Structured Data from Queries"Lecture Notes in Artificial Intelligence. 2225. 315-331 (2001)
H.Arimura 等人:“从查询中高效学习半结构化数据”人工智能讲座笔记。
DOI: --
发表时间:
期刊:
影响因子: --
作者: []
通讯作者:
39
    Next-Generation Semi-structured Data Mining for Large-Scale Knowledge Base Formation
    • 批准号:
      20240014
    • 项目类别:
      Grant-in-Aid for Scientific Research (A)
    • 资助金额:
      $32.28万
    • 财政年份:
      2008
    • 负责人:
      ARIMURA Hiroki
    • 依托单位:
    海外基金