Development of Language Information System and Systematic Extraction of Latent Knowledge, Based on Graph Computation of Semantic Network
Development of Language Information System and Systematic Extraction of Latent Knowledge, Based on Graph Computation of Semantic Network
批准号:
18500192
负责人:
AKAMA Hiroyuki
金额:
$2.05万
依托单位国家:
日本
项目类别:
Grant-in-Aid for Scientific Research (C)
财政年份:
2006
资助国家:
日本
项目状态:
已结题
起止时间:
2006 至 2007
中文摘要
我们提出了一种新的方法来解决使用MCL处理文档和语料库时观察到的聚类大小不平衡问题。分支MCL(BMCL)或潜在邻接矩阵可以将过度包容的马尔可夫簇(核心簇)调整为适当的子集。该方法被应用于从Gakken的日语大词典(GLDJ)的大规模语料库构建的语义网络,该语料库涵盖了100,000个单词、定义、实例和语法解释。这些技术的有效性目前正在通过为GLDJ创建聚类语义网络来检验,作为图聚类在人文领域的应用,我们从一些历史文献或小说的词汇共现数据中构建了语义网络:用来衡量两位当代思想家卡巴尼斯和梅斯默的书之间的思维相似性;用圣埃克苏佩里的著名小说《小王子》客观地提出了一种适用于他的谜语用法的词义消解方法。在这项研究中,我们提出了一种新的窗口方法,称为增量推进窗口(TAW),它生成的共现单词对可以用作增量路由算法的输入,作为基于关键字的聚类的替代方案。使用加权曲率、模数Q和F度量等指标对MCL应用于共现和/或邻接数据矩阵的结果进行了评估。
英文摘要
We developed a new method of solving the cluster-size imbalance problem observed when documents and corpora are processed with MCL. The Branching MCL (BMCL) or the latent adjacency matrix can resize overly inclusive Markov clusters (core clusters) into appropriate subsets. This method is applied to a semantic network built from the large-scale corpus of Gakken's Large Dictionary of Japanese (GLDJ), covering 100,000 words, definitions, examples, and grammatical explanations. The effectiveness of these techniques is currently being tested by creating a clustered semantic network for the GLDJ.As the applications of the graph clustering to the field of Humanities, we made the semantic networks from the lexical co-occurrence data of some historical documents or novels : the books of two contemporary thinkers, Cabanis and Mesmer, to measure the similarity of thinking between them ; the very famous novel of Saint-Exupery, "Le petit prince" to objectively propose a method of word sense disambiguation applicable to his enigmatic word usage. In this study we proposed as an alternative to the keyword-based clustering a new windowing method called Incrementally Advancing Window (TAW) that generates co-occurring word pairs that can be used as inputs to the Incremental Routing Algorithm. The results of the MCL applied to co-occurrence and/or adjacency data matrices were evaluated by using the indexes as weighted curvature, modularity Q and F measure.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
テキスト分析における2部グラフクラスタリングの可能性
二分图聚类在文本分析中的潜力
DOI:
--
发表时间:
2006
期刊:
情報処理学会研究報告 NL-174
影响因子:
--
作者:
[赤間啓之, 三宅真紀, 鄭在玲]
通讯作者:
鄭在玲
グラフクラスタリングを用いた文献解析の諸技法に関して-カバニスとメスメルのテキストを例に
关于使用图聚类进行文献分析的各种技术——以Cabanis和Mesmer的文本为例
DOI:
--
发表时间:
2007
期刊:
影响因子:
--
作者:
[Hiroyuki, Akama, 赤間啓之]
通讯作者:
赤間啓之
L'elaboration d'un reseau semantique par le raffinement du Markov Clustering -A partir des donnees lexicales du roman de Saint-Exupery, 《Lepetit prince》
马尔可夫聚类提纯语义的详细阐述 -圣埃克苏佩里罗马词典的一部分,《Lepetit Prince》
DOI:
--
发表时间:
2008
期刊:
Actes des 9es Journees internationales d'Analyse Statistique des Donnees Textuelles, Lyon, Presses Universitaires de Lyon 1
影响因子:
--
作者:
[Hiroyuki Akama, Maki Miyake, Jaeyoung Jung]
通讯作者:
Jaeyoung Jung
Development of a Web-based Composition Support System - Using Graph Clustering Methodologies Applied to an Associative Concepts Dictionary
基于网络的作文支持系统的开发 - 使用应用于关联概念词典的图聚类方法
DOI:
--
发表时间:
2006
期刊:
ICALT-2006
影响因子:
--
作者:
[Jaeyoung Jung, Maki Miyake, Nobuyasu Makoshi, Hiroyuki Akama]
通讯作者:
Hiroyuki Akama
DOI:
--
发表时间:
2008
期刊:
The Mathematical Intelligencer
影响因子:
--
作者:
[Hiroyuki Akama;Maki Miyake;Jaeyoung Jung]
通讯作者:
Hiroyuki Akama;Maki Miyake;Jaeyoung Jung
共 27 条
Research attempt of computational graph-based neurolinguistics integrating brain fMRI, machine learning and complex networks
-
批准号:23500171
-
项目类别:Grant-in-Aid for Scientific Research (C)
-
资助金额:$2.75万
-
财政年份:2011
-
负责人:AKAMA Hiroyuki
-
依托单位:
Study of cultural history concerning the French Ideology
-
批准号:09610501
-
项目类别:Grant-in-Aid for Scientific Research (C)
-
资助金额:$1.09万
-
财政年份:1997
-
负责人:AKAMA Hiroyuki
-
依托单位:
海外基金