Clustering and Categorization of Brazilian Portuguese Legal Documents
Clustering and Categorization of Brazilian Portuguese Legal Documents
复制标题
巴西葡萄牙语法律文件的聚类和分类
DOI:
10.1007/978-3-642-28885-2_31
复制
发表时间:
2012
期刊:
影响因子:
--
通讯作者:
Vera Lúcia Strube de Lima
中科院分区:
文献类型:
--
作者:
Luis Otávio de Colla Furquim;Vera Lúcia Strube de Lima
This study explores the use of machine learning in case law search in electronic trials. We clustered case law documents, automatically generating classes to a categorizer. These classes are used when a user uploads new documents to an electronic trial. We selected the algorithm TClus, created by Aggarwal, Gates and Yu, removing its document/group discarding features and adding a cluster division feature. We introduced a new paradigm “bag of terms and law references” instead of “bag of words” by generating attributes using a law domain thesaurus to detect legal terms and using regular expressions to detect law references. We clustered a case law corpus. The results were evaluated with the Relative Hardness Measure (RH) and the-Measure (RHO). The results were tested both with Wilcoxon’s Signed-ranks Test and Count of Wins and Losses Test to determine their significance. The categorization results were evaluated by human specialists. We compared true/false positives against document similarity with the centroid, cluster size, quantity and type of the attributes in the centroids and cluster cohesion. The article also discusses attribute generation and its implications to the classification results.