Collaborative Knowledge Discovery in Digital Government Data Using Distributed Higher-Order Text Mining
Collaborative Knowledge Discovery in Digital Government Data Using Distributed Higher-Order Text Mining
批准号:
0703698
负责人:
William Pottenger
金额:
$0.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2006
资助国家:
美国
项目状态:
已结题
起止时间:
2006-10-01 至 2009-12-31
中文摘要
分布式数据源中文本数据的迅速增长,加上创建和维护中央存储库所涉及的障碍,激发了对有效的分布式信息提取和挖掘技术的需求。给定个人的不同类型的记录可能存在于不同的数据库中——这是一种数据碎片。然而,即使有了标准,自动集成模式的能力仍然是一个开放的研究问题。一个相关的问题是,当前用于挖掘分布式数据的关联规则挖掘(Association Rule Mining, ARM)算法只有在已知所有数据库的全局模式时才能够挖掘数据(无论是垂直碎片还是水平碎片)。在从分布式文本数据提取信息的情况下,没有预先存在的全局模式可用。这是因为在不同的文档中提取的实体不同——新的输入文本可能包含以前未见过的实体。因此,不能假设固定的全局模式,也不能使用现有的算法。这项工作描述了一个分布式的高阶文本挖掘框架,它既不需要全局模式的知识,也不需要将模式集成作为挖掘规则的前提。这个框架被称为D-HOTM,它提取实体,并根据由公共键链接的记录中实体之间的高阶关联发现规则。实体提取是基于使用半监督主动学习算法学习的信息提取规则。学习到的规则被应用于从描述的文本数据中自动提取实体,例如,犯罪手法。提取的实体存储在本地关系数据库中,使用D-HOTM分布式关联规则挖掘算法进行挖掘。这项工作的更广泛影响在于与当地执法部门和医疗保健提供者合作,部署现场试验台,通过挖掘报告和确定医生最佳做法来解决问题。为学生提供大学前实习,并为研究生提供支持。
英文摘要
ABSTRACTNSF-0534276Pottenger, WilliamThe burgeoning amount of textual data in distributed sources combined with the obstacles involved in creating and maintaining central repositories motivates the need for effective distributed information extraction and mining techniques. Different kinds of records on a given individual may exist in different databases - a type of data fragmentation. Even with standards, however, the ability to integrate schemas automatically is an open research issue. A related issue is the fact that current Association Rule Mining (ARM) algorithms for mining distributed data are capable of mining data (whether vertically or horizontally fragmented) only when the global schema across all databases is known. In the case of information extracted from distributed textual data, no preexisting global schema is available. This is due to the fact that the entities extracted vary between documents - new input text can contain previously unseen entities. As a result, a fixed global schema cannot be assumed and existing algorithms cannot be employed.This effort describes a distributed higher-order text mining framework that requires neither the knowledge of the global schema nor schema integration as a precursor to mining rules. The framework, termed D-HOTM, extracts entities and discovers rules based on higher-order associations between entities in records linked by a common key. The entity extraction is based on information extraction rules learned using a semi-supervised active learning algorithm previously developed. The rules learned are applied to automatically extract entities from textual data that describe, for example, criminal modus operandi. The entities extracted are stored in local relational databases, which are mined using the D-HOTM distributed association rule mining algorithm.The broader impacts of thework lie in the collaboration with local law enforcement and healthcare providers for deploying live test beds that enable problem solving by mining reports and identificaiton of physician best practices. Pre-college internships are provided for students as well as support for graduate students.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
III: RI: Small: Efficient Privacy Methods Using Linear Programming
-
批准号:1018445
-
项目类别:Standard Grant
-
资助金额:$49.93万
-
财政年份:2010
-
负责人:William Pottenger
-
依托单位:
III: Visual Analytics for Steering Large-Scale Distributed Data Mining Applications
-
批准号:0712139
-
项目类别:Continuing Grant
-
资助金额:$44.0万
-
财政年份:2007
-
负责人:William Pottenger
-
依托单位:
Collaborative Knowledge Discovery in Digital Government Data Using Distributed Higher-Order Text Mining
-
批准号:0534276
-
项目类别:Continuing grant
-
资助金额:$0.0万
-
财政年份:2006
-
负责人:William Pottenger
-
依托单位:
Digital Government: Social Processes and Content in Intelink Online Chat Data
-
批准号:0196374
-
项目类别:Standard Grant
-
资助金额:$3.02万
-
财政年份:2001
-
负责人:William Pottenger
-
依托单位:
海外基金