An Analysis of Image Cognition Structure for Metadata Type Image Retrieval
An Analysis of Image Cognition Structure for Metadata Type Image Retrieval
复制标题
元数据型图像检索的图像认知结构分析
DOI:
--
复制
发表时间:
2001
期刊:
影响因子:
--
通讯作者:
K. Akahori
中科院分区:
文献类型:
--
作者:
T. Fukumoto;K. Akahori
Nowadays, a large amount of digital images are being stored worldwide in the Internet. As an educational means, images stored in it have big potential. And it is so rapidly expanding and becomes so complicated that the ways to retrieve images effectively are getting more difficult. As a Digital Disc Camcorder or movie delivery via the Internet, not only still images but also motion pictures will be stored and retrieved seamlessly. We considered the structure of metadata suitable for storing and retrieving seamlessly from one system both them. We examined what part in them people pay attention to and the difference between the parts. The order of keyword attached to images is proper noun, situation, event, action of main object, action without the main object or background, and impression. And the subject describes the keyword first which has an affective impact in an image. And we may say that the structure of metadata needs to be prepared for three patterns: still image, motion picture like still, and motion picture. Appling these rules, database administrators efficiently attach keywords to images. At retrieval, the system will show better results. And users store and retrieve suitable data for each media type or content. 1.Introduction Over the years large amounts of computer-aided images are being stored in the Internet owing to widely available digital recording devices. As a Digital Disc Camcorder or movie delivery via the Internet, not only still images but also motion pictures will be stored and retrieved seamlessly. There are 3 kinds of image database: feature type, sensitiveness type, and metadata type. Our concern is that at one system will seamlessly be stored and retrieved both still images and motion pictures. So feature type and sensitiveness type are not suitable for retrieving motion pictures. Because it is difficult to extract color histograms or shape from them (MPEG-7,2001). This is why we focus on metadata type in this paper. Of metadata type, first, database creators define the structure or the framework of metadata. Second, database administrators attach metadata to images according to it in the database. Third, a retriever specifies texts as a key to the database. Finally the database system searches images using the metadata which is given by the administrator and also using the texts which are keyed in by the retriever. Examples of metadata are keywords, texts, classification items and so on. By the way, we may note, in passing, the indexing of image. In the area of art documentation, there are some trials to describe picture content (UEDA,1997), 4W method, PDL method, and Iconography. But they mainly depend on the painter’s intention or bibliography of a picture itself. So they are hard to apply image retrieval. 2. The Framework of Metadata We have already dealt with the criteria of metadata/keyword (FUKUMOTO,2000). For storing and retrieving seamlessly at one system both still images and motion pictures, we are now concerned with the structure of metadata suitable for them. For this purpose, we examine what part in motion pictures or still images people pay attention to. If approved, this will be a guide on attaching metadata for database administrators. And is it different the parts between motion pictures and still images that people pay attention? If so, the structure of metadata attached to still images and to motion pictures should be different. This leads database users to store and retrieve contents suitable for each media type. In the next section the method of our experiment. In Section 4 the result is discussed. In Section 5 our conclusion is presented and the future work is discussed 3. The Procedure of Experiment Subjects of the experiment are 28 undergraduate students. All are accustomed to search engine in the Internet. First of all, We divided 28 students randomly into two groups; Group A and Group B. Group A were showed the motion pictures first and the still images subsequently. Group B were showed the still images first and motion pictures subsequently. Next we let them to show still image or motion picture. The contents are 13 patterns of various types: news, soccer, baseball, drama, sight, and animation. And we prepared two or more contents of each type (news, soccer.) The motion picture consists of one scene, from a switch point of a scene to the next one, and the still image is extracted as key frame of the motion picture. The motion picture and corresponding still image are same contents for the previous reason. We let the motion picture repeat three times and still image show as much time as the motion picture corresponding to. Then we let them give keywords to both. They gave as many keywords to them as they liked. We did not limit the number of keywords. 4.The Result of Experiment 4.1 Details and Orders of