Machine Translation and Automated Analysis of Cuneiform Languages
Machine Translation and Automated Analysis of Cuneiform Languages
批准号:
329145082
负责人:
Professor Dr. Christian Chiarcos
金额:
$0.0万
依托单位国家:
德国
项目类别:
Research Grants
财政年份:
2017
资助国家:
德国
项目状态:
已结题
起止时间:
2016-12-31 至 2020-12-31
中文摘要
古代美索不达米亚的历史和文化,第一个帝国的家园和文字的诞生地,主要是通过文学和皇家铭文。然而,占所有楔形文字文件90%以上的行政文本受到的关注要少得多:即使经过音译和数字化,大多数仍然没有翻译,因此即使是密切相关领域的学者也无法访问。但这些文本是独特的和深刻的社会历史见证,因为它们记录了早期国家经济的日常管理。由于它们数量庞大,人工翻译似乎是一项无法实现的任务。从21世纪起。仅在公元前,我们就可以访问超过67,000个数字翻译,这些数字翻译是由专家在特定的文件子集中定期制作的,但如果没有翻译,即使是亚述学其他分支的专家也难以解释。MTAAC将机器学习(ML)的最新发展与统计和神经机器翻译(MT)相结合,以促进对这些材料的分析,从而从根本上扩大其对人文和社会科学的可访问性。主要成果是一种方法,其实施以及一系列翻译和分析的文本,在开放许可下发布。除了楔形文字研究,我们还为处理不同历史语言学中的大量可比数据集树立了榜样。由于文本如此之多,我们用自动化解决方案来补充人力劳动。自然语言处理的统计和神经方法在过去几十年中已经成熟,并得到了广泛的使用,但很少应用于主要的历史语言。我们的目标是弥合这一差距,为ML和MT在人文学科中树立榜样,并促进楔形文字语言的研究。为了提高可重用性,我们根据链接的开放数据形式主义调整和开发社区维护的规范,并提出与其他数字人文行为者(如博物馆和各种语言学门户网站)合作的最佳实践规则。 加拿大多伦多大学的PI石楠贝克领导了MTAAC语言特定方面的工作。加州大学洛杉矶分校的合作PI Robert Englund是Cuneiform数字图书馆计划的负责人,负责数据管理和托管。德国歌德大学法兰克福的联合PI Christian Chiarcos负责ML、MT和数据集成。方法是协作制定的。MTAAC提供了对早期写作的高度代表性语料库的统一访问,并将采用MT和ML来促进其上下文敏感的语义解释。该项目将促进各学科研究人员之间前所未有的学术合作。因此,将使网络公众能够使用传播渠道,了解已消亡数千年的文明遗产,从而有助于更深入地欣赏和了解现代文化及其历史根源。
英文摘要
History and culture of ancient Mesopotamia, home of the first empires and birthplace of writing, are mostly known through literary and royal inscriptions. Yet, administrative texts, that make up well over 90% of all cuneiform documents, have received much less attention: Even when transliterated and digitized, most remain untranslated, and therefore inaccessible to scholars in even closely related fields. But these texts are unique and deeply insightful socio-historical witnesses, as they document the day-to-day management of early state economies. Because of their vast numbers, their human translation appears to be an unachievable task. From the 21st c. BC alone, we have access to more than 67,000 digital transcriptions as routinely produced by specialists in a particular subset of documents, but without translation difficult to interpret even by specialists in other branches of Assyriology. MTAAC combines recent developments in machine learning (ML) with statistical and neural machine translation (MT) to facilitate the analysis of this material, thereby fundamentally expanding its accessibility to the Humanities and Social Sciences.Main outcome is a methodology, its implementation, and a body of translated and analyzed texts, released under open licenses. Beyond cuneiform studies, we set an example for processing a host of comparable datasets in different historical philologies. Because the texts are so numerous, we supplement human labor with automated solutions. Statistical and neural approaches to Natural Language Processing have been maturing in the last decades, and enjoy wide usage, but have rarely been applied to even major historical languages. We aim to bridge this gap, set an example for ML and MT in the Humanities, and facilitate studies of cuneiform languages. To increase re-usability, we adapt and develop community-maintained specifications based on linked open data formalisms, and propose rules of best practice for collaboration with other digital humanities actors such as museums, and portals for various strands of philology. PI Heather Baker, University of Toronto, Canada, leads the work on language specific aspects in MTAAC. Co-PI Robert Englund, UCLA, director of the Cuneiform Digital Library Initiative, is in charge of data management and hosting. Co-PI Christian Chiarcos, Goethe University Frankfurt, Germany, is responsible ML, MT and data integration. Methodologies are developed collaboratively. MTAAC provides unified access to a highly representative corpus of early writing, and will employ MT and ML to facilitate its context-sensitive semantic interpretation. The project will foster an unprecedented scholarly cooperation among researchers in a variety of disciplines. As a result, lines of communication to the heritage of civilizations dead for many millennia will be made accessible to the networked public, contributing to a deeper appreciation and understanding of modern culture and its historical roots.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Continuation of the constitution of the "Virtual Library for General Linguistics and ComparativeLanguage Studies" in the context of the Special Subject Collection Linguistics 7.11 funded by theGerman Research Society
-
批准号:214512695
-
项目类别:Acquisition and Provision (Scientific Library Services and Information Systems)
-
资助金额:$0.0万
-
财政年份:2011
-
负责人:Professor Dr. Christian Chiarcos
-
依托单位:
海外基金