Identification of Ancient Greek Papyrus Fragments Using Genetic Sequence Alignment Algorithms

Identification of Ancient Greek Papyrus Fragments Using Genetic Sequence Alignment Algorithms
复制标题

使用遗传序列比对算法识别古希腊纸莎草碎片

DOI:
10.1109/escience.2014.14
复制
发表时间:
2014
期刊:
2014 IEEE 10th International Conference on e-Science
影响因子:
--
通讯作者:
Haoyu Yu
Haoyu Yu
中科院分区:
--
文献类型:
--
作者:
Alex C. Williams;Hyrum D. Carroll;John F. Wallin;James H. Brusuelas;L. Fortson;A. Lamblin;Haoyu Yu

文献摘要

被引文献

相似文献

纸莎草学家分析,转录和编辑纸莎草碎片,以便通过更好地理解古代世界的语言学,文化和文学来丰富现代生活。他们的共同任务之一是将未知的碎片与已知的手稿相匹配。当碎片被损坏并且仅包含有限的信息(例如,由于变质)。在过去的100年里,从埃及村庄Oxyrhynchus发现的50多万个碎片中只有大约10%被编辑过。我们不知道可能会发现什么新的古代文献,也不知道从中可以学到什么,但使用目前的鉴定方法,这一过程将需要1000多年。在计算生物学中,用一组已知的文本序列来识别一个匿名的字符串是普遍存在的。基因通常由一系列连续的字符表示,每个字符代表一个氨基酸。通过查找匿名序列和已知序列之间共享的多字母模式来推断关系。这个过程通常被称为基因序列比对。在本文中,我们介绍了一种新的方法,该方法使用现代基因序列比对算法作为识别古希腊文本片段的方法。该应用程序将为纸莎草学家和其他人文专业人员提供快速识别严重受损文本的能力。这种方法利用了一种新形式的非上下文、多行希腊语文本识别,可以大大加快繁琐的转录和识别任务。
Papyrologists analyze, transcribe, and edit papyrus fragments in order to enrich modern lives by better understanding the linguistics, culture, and literature of the ancient world. One of their common tasks is to match an unknown fragment to a known manuscript. This is especially challenging when the fragments are damaged and contain only limited information (e.g., due to deterioration). In the last 100 years, only about 10% of the more than 500,000 fragments recovered from the Egyptian village of Oxyrhynchus have been edited. We do not know what new ancient texts might be found and what can be learned from them, but using current methods of identification this process will take in excess of 1000 years. The identification of an anonymous string of characters with a collection of known text sequences is ubiquitous in computational biology. Genes are often represented by a sequence of continuous characters, each of which denotes an amino acid. Relationships are inferred by finding multi-letter patterns shared between the anonymous sequence and a known sequence. This process is commonly referred to as genetic sequence alignment. In this paper, we introduce a novel methodology that uses modern genetic sequence alignment algorithms as a method for identifying Ancient Greek text fragments. This application will offer papyrologists and other professionals in the humanities the ability to rapidly identify severely damaged texts. This approach leverages a new form of non-contextual, multi-line text identification for the Greek language that can greatly accelerate the tedious task of transcription and identification.