The Archaeotools project: faceted classification and natural language processing in an archaeological context

The Archaeotools project: faceted classification and natural language processing in an archaeological context
复制标题

DOI:
10.1098/rsta.2009.0038
复制
发表时间:
2009-06-28
影响因子:
5
通讯作者:
Zhang, Z.
Zhang, Z.
中科院分区:
综合性期刊2区
文献类型:
--
作者:
Jeffrey, S.;Richards, J.;Zhang, Z.

文献摘要

被引文献

相似文献

本文介绍了考古学中的一个重要的电子科学项目“考古工具”。该项目的目的是使用分面分类和自然语言处理为考古研究创建先进的基础设施。该项目旨在将涉及英国考古遗址和古迹的超过1x10(6)个结构化数据库记录与从半结构化灰色文献报告和非结构化古物期刊账户中提取的信息整合在一个单一的浏览器界面中。该项目阐明了目前存在于国家和地方古迹目录中的词汇控制和标准化的可变水平。尽管如此,它已经表明,相对定义良好的本体和叙词表,存在于考古学意味着一个高水平的成功,可以实现使用信息提取技术。这对于解锁和获取灰色文献和古董商账户中的信息具有巨大的潜力,并为相关学科提供了经验教训。
This paper describes 'Archaeotools', a major e-Science project in archaeology. The aim of the project is to use faceted classification and natural language processing to create an advanced infrastructure for archaeological research. The project aims to integrate over 1 x 10(6) structured database records referring to archaeological sites and monuments in the UK, with information extracted from semi-structured grey literature reports, and unstructured antiquarian journal accounts, in a single faceted browser interface. The project has illuminated the variable level of vocabulary control and standardization that currently exists within national and local monument inventories. Nonetheless, it has demonstrated that the relatively well-defined ontologies and thesauri that exist in archaeology mean that a high level of success can be achieved using information extraction techniques. This has great potential for unlocking and making accessible the information held in grey literature and antiquarian accounts, and has lessons for allied disciplines.