课题基金 / 基金详情

Development of efficient knowledge discovery systems for large semistructured data

Development of efficient knowledge discovery systems for large semistructured data
开发针对大型半结构化数据的高效知识发现系统
批准号:
17200011
负责人:
OKAMOTO Seishi
金额:
$29.62万
依托单位:
依托单位国家:
日本
项目类别:
Grant-in-Aid for Scientific Research (A)
财政年份:
2005
资助国家:
日本
项目状态:
已结题
起止时间:
2005 至 2007

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
By the rapid progress of Internet and Web service technologies, a new kind of massive data called semistructured data emerged, where a semistructured data is a collection of weakly structured electronic data such as Web pages and XML documents. In this research project, we studied efficient knowledge discovery systems for large semistructured data.First we studied theoretical foundations of learning and discovery for semistructured data. One of our main contributions is on kernels for trees. We introduced a new kernel function for labeled ordered trees and showed a hardness result in designing tree kernels for more general labeled trees(JSAI Best Paper Award in 2006). Another important one is on episode mining. We showed that an episode is parallel-free if and only if it is serially constructive.Next, we studied practical processing methods for semistructured data such as pattern matching, text compression, and index structures. Main contributions are as follows We devised efficient matching algorithms for path patterns based on the one-way sequential processing. These algorithms run 2 - 6 times faster and 6 times space-efficient in comparison with XMLTK. We also proposed an efficient index structure for the fast reachability test on directed graphs and implemented it(DEWS2007 BestPaper Award). Furthermore, we developed a new compressed pattern matching(CPM) algorithm that improves both the compression ratio and the search time ratio in comparison with a BPE type CPM algorithm.Finally, we applied the theoretical and practical results in this project to knowledge discovery systems. We demonstrated that these applications work effectively in various areas such as bioinformatics, pharmacy, music, traffic, and security.
期刊论文(153)
专著(0)
科研奖励(0)
会议论文
Mining Frequent Elliptic Episodes from Event Sequence
从事件序列中挖掘频繁的椭圆情节
DOI: --
发表时间: 2007
期刊: Proc.the 5th Workshop on Learning with Logic and Logics for Learning
影响因子: --
作者: [Takashi Katoh, 他1名]
通讯作者: 他1名
Reducing Trials by Thinning-Out in Skill Discovery
通过技能发现中的稀疏化来减少试验
DOI: --
发表时间: 2007
期刊: Lecture Notes in Computer Science(Proc.the 10th International Conference on Discovery Science) 4755
影响因子: --
作者: [Hayato Kobayashi, 他3名]
通讯作者: 他3名
An Efficient Algorithm for Complex Pattern Matching over Continuous Data Streams Based on Bit-Parallel Method
基于位并行方法的连续数据流复杂模式匹配的高效算法
DOI: --
发表时间: 2007
期刊: Proc.3^<rd> IEEE International Workshop on Databases for Next-Generation Researchers
影响因子: --
作者: [Tomoya Saito, 他2名]
通讯作者: 他2名
連続データストリームに対するビット並列手法を用いた高度な時系列パターン照合
使用位并行技术进行连续数据流的高级时间序列模式匹配
DOI: --
发表时间: 2007
期刊: 第18回データ工学ワークショップ(DEWS2007)
影响因子: --
作者: [斉藤 智哉, 他2名]
通讯作者: 他2名
115
    国内基金
    海外基金
    基于集成学习的分布式XML数据流的挖掘模型与概念漂移挖掘方法研究
    • 批准号:
      61773415
    • 项目类别:
      面上项目
    • 资助金额:
      64.0万元
    • 批准年份:
      2017
    • 负责人:
      毛国君
    • 依托单位:
    海量不确定XML数据查询关键技术研究
    • 批准号:
      61602130
    • 项目类别:
      青年科学基金项目
    • 资助金额:
      20.0万元
    • 批准年份:
      2016
    • 负责人:
      刘健
    • 依托单位:
    高扩展性XML关键字查询处理技术
    • 批准号:
      61572421
    • 项目类别:
      面上项目
    • 资助金额:
      66.0万元
    • 批准年份:
      2015
    • 负责人:
      陈子阳
    • 依托单位:
    基于事前约束的XML关键字查询处理技术
    • 批准号:
      61472339
    • 项目类别:
      面上项目
    • 资助金额:
      80.0万元
    • 批准年份:
      2014
    • 负责人:
      周军锋
    • 依托单位: