Named Entity Recognition and Resolution in Legal Text

Named Entity Recognition and Resolution in Legal Text
复制标题

法律文本中命名实体的识别和解析

DOI:
10.1007/978-3-642-12837-0_2
复制
发表时间:
2010
期刊:
--
影响因子:
--
通讯作者:
Ramdev Wudali
Ramdev Wudali
中科院分区:
--
文献类型:
--
作者:
Christopher C. Dozier;R. Kondadadi;M. Light;Arun Vachher;S. Veeramachaneni;Ramdev Wudali

文献摘要

被引文献

相似文献

文本中的命名实体是文本中使用专有名词明确提及的人、地点、公司等。在文本中发现命名实体并将其分类为语义类型的过程称为命名实体识别。命名实体的解析是将文本中提到的名称与预先存在的数据库条目联系起来的过程。这是提到类似于一个真实的世界实体的东西的理由。例如,提到一个名叫Mary Smith的法官,可能会被解析为一个数据库条目,该条目对应于一个特定州的一个特定地区的一个特定法官。这种命名实体的识别和解析可以通过多种方式来利用,包括提供到存储的关于特定法官的信息的超文本链接:他们的教育、谁任命他们、他们的其他案件意见等。本文讨论了法律的文件中的命名实体识别和解析,如美国判例法、证词、诉状和其他审判文件。实体的类型包括法官、律师、公司、司法管辖区和法院。我们概述了命名实体识别的三种方法,查找,上下文规则和统计模型。然后,我们描述了一个实际的系统,发现命名实体的法律的文本和评估其准确性。同样,对于分辨率,我们讨论了我们的分块技术,我们的分辨率特征,以及我们用于最终匹配的监督和半监督机器学习技术。
Named entities in text are persons, places, companies, etc. that are explicitly mentioned in text using proper nouns. The process of finding named entities in a text and classifying them to a semantic type, is called named entity recognition. Resolution of named entities is the process of linking a mention of a name in text to a pre-existing database entry. This grounds the mention in something analogous to a real world entity. For example, a mention of a judge namedMary Smithmight be resolved to a database entry for a specific judge of a specific district of a specific state. This recognition and resolution of named entities can be leveraged in a number of ways including providing hypertext links to information stored about a particular judge: their education, who appointed them, their other case opinions, etc.This paper discusses named entity recognition and resolution in legal documents such as US case law, depositions, and pleadings and other trial documents. The types of entities include judges, attorneys, companies, jurisdictions, and courts.We outline three methods for named entity recognition, lookup, context rules, and statistical models. We then describe an actual system for finding named entities in legal text and evaluate its accuracy. Similarly, for resolution, we discuss our blocking techniques, our resolution features, and the supervised and semi-supervised machine learning techniques we employ for the final matching.