课题基金 / 基金详情

Detecting relevant segment of text in legal domain

Detecting relevant segment of text in legal domain
检测法律领域中的相关文本片段
批准号:
499514-2016
负责人:
Makrehchi, Masoud
金额:
$1.82万
依托单位国家:
加拿大
项目类别:
Engage Grants Program
财政年份:
2016
资助国家:
加拿大
项目状态:
已结题
起止时间:
2016-01-01 至 2017-12-31

项目摘要

项目成果

Makrehchi, Masoud的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
The goal of the research is to investigate, design, and implement algorithms to detect (or recognize) and extract the relevant segment of text, predict and recognize legal entities and context, and finally generate an appropriate metadata to be stored in a structured database. The database can be utilized in several scenarios from the user query to legal research by law practitioners. The notion of "relevant segment" is defined as a contiguous piece of a text which is relevant to the question of interest (or simply query). Relevance can be measured by different methods depending how relevance is being interpreted. If we are looking for the name of a judge in a legal document, we can use a wide range of information extraction (IE) tools. IE takes advantage of a broad spectrum of techniques from image segmentation, when the image of the document is available and a relevant segment is highly expected in a specific zone, to Conditional Random Fields (CRF) and Markov Models to Machine Learning and classification. While the structured pieces of information such as entities can be extracted using IE techniques, for deeper, ambiguous, and conceptual components of a legal document such as the type of damage or the judge's decision and case outcome, we need to develop a supervised machine learning algorithms beyond IE techniques. This problem is neither a traditional IE problem nor a text classification. To solve this problem, a legal document is partitioned into conceptually-related segments such as header, case, citations, damages, decision, and so on. This step is called zoning and can be performed using supervised or unsupervised learning methods. Some zones such as headers are expected to appear in the very first section of the document and so they can be detected by unsupervised techniques. On the other hand, there are other components such as "damages" which may appear in any part of the documents and needs a supervised model using either lexicon-based or manually-labeled grand truth or both.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Algorithms and applications of Link Mining: Making Sense of Network Data
Algorithms and applications of Link Mining: Making Sense of Network Data
Towards Predicting Socio-economic Systems by Mining Social Media Data
Towards Predicting Socio-economic Systems by Mining Social Media Data
海外基金