EAGER: Automatic Document and Record Disposition and Retention
EAGER: Automatic Document and Record Disposition and Retention
批准号:
1143921
负责人:
C. Lee Giles
金额:
$20.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2011
资助国家:
美国
项目状态:
已结题
起止时间:
2011-08-01 至 2015-07-31
中文摘要
记录和文件保留(文件处理)已成为组织和个人的一个严重问题,因为现在创建的大多数文件都是数字化的。数字文档既有问题,也有优势。数字文档很容易进行版本化、复制和传播。因此,在许多位置可能存在重要或相关文件的几个类似副本或版本。文档或记录处理可以应用于个人、组织和领域(如法律、科学、政策等),或个人、组织和领域(如法律、科学、政策等)所需要的。用于长期有效的信息管理。这个问题是史诗般的规模,并正在成为世界各地的组织和个人的一个主要问题,因为法律或组织或系统中的实际限制要求有效的记录处理。这个探索性项目研究了基于文本检查、挖掘和搜索算法的可能的自动文档处理方法。挑战在于寻找可伸缩的、可适应的算法,这些算法可以在几个(如果不是所有的)应用程序领域中使用。此外,用户的可变性带来了许多问题。处置方法或程序可根据用户、组织和领域(例如,法律、健康记录等)而有所不同。本项目中探讨的方法将机器学习方法应用于并扩展到这些问题,因为这些方法适应数据、领域和领域的可变性。使用这些方法,自动处理方法可以很容易地应用于这些不同的领域,如科学、电子邮件和法律记录。这项研究为各种领域的自适应方法在适用性、性能和可扩展性方面奠定了基础。这一概念验证项目最初侧重于可公开获得的安然电子邮件数据集,可用来证明该方法的可行性,因为电子邮件可被视为文件处理的特例。如果成功,还将探索其他处置领域,如科学和政府数据。这项工作将展示开发和应用机器学习方法到一个重要和多样化的问题领域的可行性。这一探索性项目的结果,以及从大规模文件搜索中使用的方法收集的见解,预计将产生关于我们如何更好地管理我们的数字过去和迅速扩展的数字未来的理解。这一结果有望将这一重要问题介绍给其他研究人员和文档处理专业人员,并导致与业界的合作。数据和研究成果将通过一个公开的网站(http://clgiles.ist.psu.edu/disposeseer/))提供,研究论文将在适当的地点发表和展示。该项目为研究生和本科生提供研究经验。
英文摘要
Record and document retention (document disposition) has become a serious problem for both organizations and individuals since most documents now created are digital. Digital documents offer both problems and advantages. Digital documents are easily versioned, copied and disseminated. Thus, there can be several similar copies or versions of important or relevant documents in many locations. Document or record disposition can be applied to or is needed by individuals, organizations and domains (such as law, science, policy, etc.) for effective information management over long periods of time. This problem is of epic proportions and is becoming a major problem in organizations and for individuals throughout the world where effective record disposition is either required by law or by the organization or by practical limitations in systems.This exploratory project investigates possible automatic document disposition methods based on algorithms for text inspection, mining, and search. The challenges lie in finding scalable, adaptable algorithms that can be used in several if not all application domains. In addition, variability in users presents many problems. A disposition method or procedure may vary depending on the user, organization and domain (e.g., law, health records, etc.). The approach explored in this project applies and extends machine learning methods to these problems since these methods adapt to variability in data, areas and domains. Using such approaches, automated disposition methods can be readily applied to these different areas such as science, email and legal records. This research lays the groundwork for adaptive methods for a variety of domains in terms of applicability, performance and scalability. this proof-of-concept project initially focuses on the Enron email data set that is publicly available and is be used to demonstrate the feasibility of the approach since email can be considered a special case of document disposition. If successful, other disposition domains such as science and government data will be explored. This work will show the viability of developing and applying machine learning methods to an important and diverse problem domain. The results from this exploratory project together with insights gathered from methods used in large scale document search are expected to yield understanding as to how we can better manage our digital past and the rapidly expanding digital future. The results are expected to introduce this important problem to other researchers and document disposition professionals and lead to collaborations with industry. Data and research results will be made available through a publicly available website (http://clgiles.ist.psu.edu/disposeseer/) and research papers will be published and presented in appropriate venues. The project provides research experience for graduate and undergraduate students.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CRI: CI-SUSTAIN: Collaborative Research: CiteSeerX: Toward Sustainable Support of Scholarly Big Data
-
批准号:1823288
-
项目类别:Standard Grant
-
资助金额:$77.0万
-
财政年份:2018
-
负责人:C. Lee Giles
-
依托单位:
III: Small: Collaborative Research: Keyphrase Extraction in Document Networks
-
批准号:1422951
-
项目类别:Continuing Grant
-
资助金额:$17.5万
-
财政年份:2014
-
负责人:C. Lee Giles
-
依托单位:
Collaborative Research: STEM Workforce Training: A Quasi-Experimental Approach Using the Effects of Research Funding
-
批准号:1348712
-
项目类别:Standard Grant
-
资助金额:$15.0万
-
财政年份:2013
-
负责人:C. Lee Giles
-
依托单位:
Collaborative Research: CI-ADDO-EN: Semantic CiteSeer X
-
批准号:0958143
-
项目类别:Continuing Grant
-
资助金额:$89.76万
-
财政年份:2010
-
负责人:C. Lee Giles
-
依托单位:
EAGER: Creating a Book Citation Index
-
批准号:1042276
-
项目类别:Standard Grant
-
资助金额:$14.97万
-
财政年份:2010
-
负责人:C. Lee Giles
-
依托单位:
CRI: Collaborative: Next Generation CiteSeer
-
批准号:0454052
-
项目类别:Continuing Grant
-
资助金额:$0.0万
-
财政年份:2005
-
负责人:C. Lee Giles
-
依托单位:
SGER: A Digital Library Archive for Computer Scientists
-
批准号:0330783
-
项目类别:Standard Grant
-
资助金额:$9.91万
-
财政年份:2003
-
负责人:C. Lee Giles
-
依托单位:
海外基金