课题基金 / 基金详情

Collaborative Knowledge Discovery in Digital Government Data Using Distributed Higher-Order Text Mining

Collaborative Knowledge Discovery in Digital Government Data Using Distributed Higher-Order Text Mining
使用分布式高阶文本挖掘的数字政府数据中的协作知识发现
批准号:
0534276
负责人:
William Pottenger
金额:
$0.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing grant
财政年份:
2006
资助国家:
美国
项目状态:
已结题
起止时间:
2006-01-01 至 2007-01-31

项目摘要

项目成果

William Pottenger的其他基金

相似基金

相关文献

中文摘要
翻译
分布式来源中迅速增长的文本数据,再加上创建和维护中央资料库所涉及的障碍,促使人们需要有效的分布式信息提取和挖掘技术。给定个人的不同类型的记录可能存在于不同的数据库中--一种数据碎片。然而,即使有了标准,自动集成模式的能力也是一个开放的研究问题。一个相关的问题是,当前用于挖掘分布式数据的关联规则挖掘(ARM)算法只有在所有数据库的全局模式已知的情况下才能够挖掘数据(无论是垂直的还是水平的碎片)。在从分布式文本数据提取信息的情况下,没有预先存在的全局模式可用。这是因为提取的实体因文档而异--新的输入文本可能包含以前看不到的实体。描述了一个分布式高阶文本挖掘框架,该框架既不需要全局模式的知识,也不需要模式集成作为挖掘规则的先导。该框架被称为D-HOTM,它提取实体并基于由公共关键字链接的记录中实体之间的高阶关联来发现规则。实体提取基于使用先前开发的半监督主动学习算法学习的信息提取规则。所学习的规则被应用于从描述例如犯罪作案手法的文本数据中自动提取实体。提取的实体存储在本地关系数据库中,使用D-HOTM分布式关联规则挖掘算法进行挖掘。这项工作的更广泛影响在于与当地执法部门和医疗保健提供商合作,部署实时试验台,通过挖掘报告和识别医生最佳实践来解决问题。大学前实习是为学生提供的,也为研究生提供支持。
英文摘要
ABSTRACTNSF-0534276Pottenger, WilliamThe burgeoning amount of textual data in distributed sources combined with the obstacles involved in creating and maintaining central repositories motivates the need for effective distributed information extraction and mining techniques. Different kinds of records on a given individual may exist in different databases - a type of data fragmentation. Even with standards, however, the ability to integrate schemas automatically is an open research issue. A related issue is the fact that current Association Rule Mining (ARM) algorithms for mining distributed data are capable of mining data (whether vertically or horizontally fragmented) only when the global schema across all databases is known. In the case of information extracted from distributed textual data, no preexisting global schema is available. This is due to the fact that the entities extracted vary between documents - new input text can contain previously unseen entities. As a result, a fixed global schema cannot be assumed and existing algorithms cannot be employed.This effort describes a distributed higher-order text mining framework that requires neither the knowledge of the global schema nor schema integration as a precursor to mining rules. The framework, termed D-HOTM, extracts entities and discovers rules based on higher-order associations between entities in records linked by a common key. The entity extraction is based on information extraction rules learned using a semi-supervised active learning algorithm previously developed. The rules learned are applied to automatically extract entities from textual data that describe, for example, criminal modus operandi. The entities extracted are stored in local relational databases, which are mined using the D-HOTM distributed association rule mining algorithm.The broader impacts of thework lie in the collaboration with local law enforcement and healthcare providers for deploying live test beds that enable problem solving by mining reports and identificaiton of physician best practices. Pre-college internships are provided for students as well as support for graduate students.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
III: RI: Small: Efficient Privacy Methods Using Linear Programming
  • 批准号:
    1018445
  • 项目类别:
    Standard Grant
  • 资助金额:
    $49.93万
  • 财政年份:
    2010
  • 负责人:
    William Pottenger
  • 依托单位:
III: Visual Analytics for Steering Large-Scale Distributed Data Mining Applications
  • 批准号:
    0712139
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $44.0万
  • 财政年份:
    2007
  • 负责人:
    William Pottenger
  • 依托单位:
Collaborative Knowledge Discovery in Digital Government Data Using Distributed Higher-Order Text Mining
  • 批准号:
    0703698
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $0.0万
  • 财政年份:
    2006
  • 负责人:
    William Pottenger
  • 依托单位:
Digital Government: Social Processes and Content in Intelink Online Chat Data
  • 批准号:
    0196374
  • 项目类别:
    Standard Grant
  • 资助金额:
    $3.02万
  • 财政年份:
    2001
  • 负责人:
    William Pottenger
  • 依托单位:
海外基金