Development of Efficient Data Mining Systems for Large Semi-Structured Text Data
Development of Efficient Data Mining Systems for Large Semi-Structured Text Data
批准号:
11558040
负责人:
ARIMURA Hiroki
金额:
$6.27万
依托单位:
依托单位国家:
日本
项目类别:
Grant-in-Aid for Scientific Research (B)
财政年份:
1999
资助国家:
日本
项目状态:
已结题
起止时间:
1999 至 2001
中文摘要
这个研究项目的目标是设计一个高效的半自动工具,支持人类从大型非结构化和半结构化的文本数据中发现。为了实现这一目标,我们从以下三个方向进行了研究.文本挖掘的核心过程是模式发现。我们采用了优化模式发现的框架,并开发了高效和强大的文本挖掘算法,发现简单的组合模式,从大型非结构化文本。为了实现这些算法,我们开发了一个适合于文本挖掘的基于后缀数组的文本索引结构。基于这些技术,我们实现了一个原型系统,并对Web数据进行了计算机实验.文本的另一个重要技术是高效的模式匹配。作为一个理论框架,我们提出了一个统一的框架,称为拼贴系统,实现各种基于字典的压缩方法。我们开发了Knuth-Morris-Pratt型和Byer-Moore型模式匹配算法,采用这个框架。并将该框架应用于Byte-Pair-Encoding压缩方法和Sequitur中,前者产生了最快的压缩模式匹配算法.文本挖掘的最后一个过程是信息抽取。从理论的角度,我们首先形式化的信息提取问题,从半结构化数据,然后给这样的任务的能力和局限性的理论分析。然后,我们开发了有效的信息提取算法的各种类型的提取规则,包括树包装和对冲模式,并评估他们通过实验在现实生活中的半结构化数据在互联网上。
英文摘要
The goal of this research project is to devise an efficient semi-automatic tool that supports human discovery from large unstructured and semi-structured text data. To achieve this goal, we studied in the following three directions.1. The central process of text mining is pattern discovery. We employed the framework of optimized pattern discovery, and developed effcient and robust text mining algorithms that find simple combinatorial patterns from large unstructured texts. To implement these algorithms, we developed a text index structure based on the suffix arrays suitable for text mining. Based on these technologies, we implemented a prototype system and run computer experiments on Web data.2. Another important technology for text is efficient pattern matching. As a theoretical framework, we proposed a unified framework, called Collage system, for realizing various dictionary-based compression methods. We developed both Knuth-Morris-Pratt type and Byer-Moore type pattern matching algorithms employing this framework. We also applied this framework to Byte-Pair-Encoding compression method and Sequitur, the former of which yields the fastest compressed pattern matching algorithm.3. Final process of text mining is information extraction. From theoretical point of view, we first formalize the information extraction problem from semi-structured data, and then gave theoretical analysis of the power and the limitation of such tasks. Then, we developed efficient information extraction algorithms for various types of extraction rules including tree wrappers and hedge patterns and evaluate them through experiments on real-life semi-structured data on the internet.
期刊论文(104)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
H. Hori et al.: "Fragmentary Pattern Matching : Complexity, Algorithms and Applications for Analyzing Classic Literary Works"Proc. 12th Annual International Symposium on Algorithms and Computation (ISAAC'01). 719-730 (2001)
H. Hori 等人:“片段模式匹配:分析经典文学作品的复杂性、算法和应用”Proc。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
K. Tamari et al: "Discovering Poetic Allusion in Anthologies of Classical Japanese Poems"Proc. 2nd Int. Conf. on Discovery Science. LNAI1721. 128-138 (1999)
K. Tamari 等:“在日本古典诗歌选集中发现诗意典故”Proc。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
M.Takeda et al.: "Mining from Literary Texts : Pattern Discovery and Similarity Computation"Lecture Notes in Computer Science. 2281. 520-533 (2002)
M.Takeda 等人:“从文学文本中挖掘:模式发现和相似性计算”计算机科学讲义。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
H.Arimura et al.: "Efficient Learning of Semi-Structured Data from Queries"Lecture Notes in Artificial Intelligence. 2225. 315-331 (2001)
H.Arimura 等人:“从查询中高效学习半结构化数据”人工智能讲座笔记。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
Tetsuya Nasukawa et al.: "Base Technology for Text Mining"Journal of Japanese Society for Artificial Intelligence. 16(2). 201-211 (2001)
那须川哲也等:《文本挖掘的基础技术》日本人工智能学会期刊。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
共 39 条
Next-Generation Semi-structured Data Mining for Large-Scale Knowledge Base Formation
-
批准号:20240014
-
项目类别:Grant-in-Aid for Scientific Research (A)
-
资助金额:$32.28万
-
财政年份:2008
-
负责人:ARIMURA Hiroki
-
依托单位:
海外基金