Advanced hierarchical information retrieval system using structure of contents, indexes and full-texts in Japanese text.
Advanced hierarchical information retrieval system using structure of contents, indexes and full-texts in Japanese text.
批准号:
06680385
负责人:
TANIGUCHI Toshio
金额:
$1.34万
依托单位:
依托单位国家:
日本
项目类别:
Grant-in-Aid for General Scientific Research (C)
财政年份:
1994
资助国家:
日本
项目状态:
已结题
起止时间:
1994 至 1995
中文摘要
本研究的目的和意义在于给未来的电子图书馆或广大的全文信息检索提供一些固定的技术指标和设计模式。在此之前,我们通过内容信息和索引构建了全文知识的框架,并将其作为教师机器来使用,并开发了检索系统将其重构为超文本。为此,我们准备并实验了以下研究过程:(1)提取全文中的术语,保持章节结构的层次结构,研究标题和内容中的层次术语,以及它们之间的差异。(2)在此基础上,定义了出现在标题、内容和索引中的术语之间的相互关系。并以此为指导,设计了一个先进的信息检索系统,并对其进行了实验。(3)从层次的角度出发,尝试对全文和…进行超文本处理最后,我们获得了20本学术书籍的机器可读全文,并以章节结构的形式将其转化为数据库。我们介绍了NAGAO博士开发的Juman:形态分析程序,并对这些文本的形式进行了分析。之后,我们排除了标准词典中的术语,并编写了程序,对复合词进行了提取,并对能够表达书本特征的术语进行了识别。内容数据的拖拽结构是由NAGAO博士设计的,可以对大部分内容进行自动拖拽。我们可以统一章节的结构和内容以及粗略和接近的区别,这是1994年尚未解决的问题。我们通过对拖拽的泛化来形式化表示起伏,并通过能够区分层次的相对起伏的检索程序来去除它们。我们利用全文名词的知识,几乎解决了粗略和接近的区别。较少
英文摘要
This research has the purpose and the meaning to give the electronic libraries or vast future full-text information retrieval some fixed technical guide lines and models of design. Before that we constructed the framework of the knowledge in the full-text by contents information and the indexes, which we used it as a teacher machine, and developed the retrieval system to reconstruct it into the hypertext. For this purpose we prepared and experimented the following research process.(1) We extracted the terms in the full-text, keeping the structure of hierarchy like chapter structure and researched the hierarchical terms in titles and contents, and differences between them.(2) According to that research we defined the mutual relatins of the terms which appeared in the titles, the contents, and the indexes. And using that as a techer we designed the advanced information retrieval system and experimented it.(3) Considering the point of hierarchy, we tried to hypertextlize the full-text and … More put that retrieval system into it.Finally we acquired the machine readable full-texts of 20 academic books, and turned them into database in the form of chapter structure. We introduced the JUMAN : morphological analysis program which Dr.NAGAO developed and analyzed the form of those texts. After that we excluded the terms in the standard dictionary and prepared the program which extracted the compound words and unidentified the terms which can express the character of the book. The tug structure of the contents data was designed by Dr.NAGAO which could do automatic tugging the most of the contents. The effects of that was equipped in Ariadne.We could unify the chapter structure of contents and difference between the rough and the close which remained unsolved 1994 year. We formalized the ups and downs by generalizing the tugging and got rid of them by the retrieval program which can distinguish the relative ups and downs of hierarchy. We almost resolved the difference of the rough and the close by using the knowledge of the nouns of the full-text. Less
期刊论文(48)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
谷口敏夫: "日本語文章における自動索引の試み" 光華女子大学研究紀要. 32. 43-63 (1994)
Toshio Taniguchi:“日文文本自动索引的尝试”Koka 女子大学研究通报 32. 43-63 (1994)。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
谷口 敏夫: "ハイパーメディア時代の図書館システム" 薬学図書館. 39. 255-263 (1994)
Toshio Taniguchi:“超媒体时代的图书馆系统”医药图书馆。39. 255-263 (1994)
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
長尾 真: "電子図書館Ariadneの開発(1)システム設計の方針" 情報管理. 38(3). 191-206 (1995)
Makoto Nagao:“电子图书馆的开发Ariadne(1)系统设计策略”信息管理191-206(1995)。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
TANIGUCHI,T.: "Library System in the age of Hypermedia" YAKUGAKU TOSYOKAN. 39. 255-263 (1994)
TANIGUCHI,T.:“超媒体时代的图书馆系统”药学图书馆。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
TANIGUCHI,T.: "Abstracting and automatic indexing in Japanese text" KOKA JYOSI-DAIGAKU KENKYU KIYO. 33. 47-81 (1995)
TANIGUCHI,T.:“日语文本中的抽象和自动索引” KOKA JYOSI-DAIGAKU KENKYU KIYO。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
共 21 条