Proceedings of the 5th ACL-HLT Workshop on Language Technology for Cultural Heritage, Social Sciences, and Humanities

Proceedings of the 5th ACL-HLT Workshop on Language Technology for Cultural Heritage, Social Sciences, and Humanities
复制标题

第五届 ACL-HLT 文化遗产、社会科学和人文语言技术研讨会论文集

DOI:
--
复制
发表时间:
2011
期刊:
影响因子:
--
通讯作者:
P. Lendvai
P. Lendvai
中科院分区:
--
文献类型:
--
作者:
Kalliopi Zervanou;P. Lendvai

文献摘要

被引文献

相似文献

LaTeCH(文化遗产、社会科学和人文科学的语言技术)年度研讨会系列旨在为研究人员提供一个论坛,这些研究人员正在研究与人文科学、社会科学和文化遗产数据有关的自然语言和信息技术应用方面。LaTeCH研讨会最初的动机是对语言技术研究和文化遗产领域应用的兴趣日益增长。然而,范围很快扩大到包括人文和社会科学。 目前网络和信息获取的发展引发了博物馆、档案馆、图书馆和其他文化遗产机构的一系列数字化努力。人文和社会科学的类似发展导致大量数据以电子格式提供,无论是数字化还是非数字化数据。数字化的下一步自然是对这些数据进行智能处理。为此,人文,社会科学和文化遗产领域吸引了越来越多的研究人员在NLP旨在开发语义丰富和信息发现和访问的方法的兴趣。语言技术通常集中在某些领域,如新闻专线。这些相当新颖的文化遗产、社会科学和人文学科领域给NLP研究带来了新的挑战,比如嘈杂的文本(例如,由于OCR问题),非标准的,或古老的语言变体(例如,历史语言、方言、语言混合使用、省略、抄写错误)、文学或比喻写作风格以及缺乏词典等知识资源。此外,通常既没有注释的领域数据,也没有手动创建它所需的资金,从而迫使研究人员调查(半)自动资源开发和领域适应方法,涉及尽可能少的人工努力。 在LaTeCH研讨会的当前版本中,我们收到了创纪录数量的提交,其中一个子集是根据全面的同行评审过程选择的。对本次LaTeCH研讨会的大多数贡献的一个中心问题是历史语言变体的语言处理问题(例如,西班牙文、捷克文、德文、斯洛文尼亚文和瑞典文)以及各自的资源开发和工具调整。在应用方面,这些贡献试图为文化遗产和人文研究人员提供语言技术解决方案,从历史学家和建筑历史学家到语言学家,文化遗产策展人,民族学家和文学评论家。分析的文本类型从全文到半结构化文本,而涉及的领域从历史文本和加密中世纪手稿的分析,到小说和童话故事以及现代学术期刊、在线博客和论坛。主题的多样性和提交数量的增加说明了对这一令人兴奋和不断扩大的研究领域的兴趣日益增长。
The LaTeCH (Language Technology for Cultural Heritage, Social Sciences, and Humanities) annual workshop series aims to provide a forum for researchers who are working on aspects of natural language and information technology applications that pertain to data from the humanities, social sciences, and cultural heritage. The LaTeCH workshops were initially motivated by the growing interest in language technology research and applications for the cultural heritage domain. The scope has soon nevertheless broadened to also include the humanities and the social sciences. Current developments in web and information access have triggered a series of digitisation efforts by museums, archives, libraries and other cultural heritage institutions. Similar developments in humanities and social sciences have resulted in large amounts of data becoming available in electronic format, either as digitised, or as born-digital data. The natural next step to digitisation is the intelligent processing of this data. To this end, the humanities, social sciences, and cultural heritage domains draw an increasing interest from researchers in NLP aiming at developing methods for semantic enrichment and information discovery and access. Language technology has been conventionally focused on certain domains, such as newswire. These fairly novel domains of cultural heritage, social sciences, and humanities entail new challenges to NLP research, such as noisy text (e.g., due to OCR problems), non-standard, or archaic language varieties (e.g., historic language, dialects, mixed use of languages, ellipsis, transcription errors), literary or figurative writing style and lack of knowledge resources, such as dictionaries. Furthermore, often neither annotated domain data is available, nor the required funds to manually create it, thus forcing researchers to investigate (semi-) automatic resource development and domain adaptation approaches involving the least possible manual effort. In the current edition of the LaTeCH workshop, we have received a record number of submissions, a subset of which has been selected based on a thorough peer-review process. A central issue for the majority of contributions to this LaTeCH workshop has been the problem of linguistic processing for historical language varieties (e.g., Spanish, Czech, German, Slovene and Swedish) and the respective resource development and tool adaptation. In terms of applications, the contributions attempt to provide language technology solutions for cultural heritage and humanities researchers ranging from historians and architecture historians to linguists, cultural heritage curators, ethnologists and literary critics. The text types targeted for analysis range from full-text to semi-structured text, while the domains addressed range from the analysis of historical text and encrypted medieval manuscripts, to novels and fairy tales and modern academic journals, online blogs and fora. The variety of topics and the increased number of submissions illustrate the growing interest in this exciting and expanding research area.