Visual descriptors

Visual descriptors

Visual descriptors describe the visual features of the contents in images or videos. They describe elementary characteristics such as the shape, the color, the texture or the motion, among others.

Introduction

As a result of the new communication technologies and the massive use of Internet in our society, the amount of audio-visual information available in digital format is increasing considerably. Therefore, it has been necessary to design some systems that allow us to describe the content of several types of multimedia information in order to search and classify them.

The audio-visual descriptors are in charge of the contents description. These descriptors have a good knowledge of the objects and events found in a video, image or audio and they allow the quick and efficient searches of the audio-visual content.

This system can be compared to the search engines for textual contents. Although it is certain, that it is relatively easy to find text with a computer, is much more difficult to find concrete audio and video parts. For instance, imagine somebody searching a scene of a happy person. The happiness is a feeling and it is not evident its shape, color and texture description in images.

The description of the audio-visual content is not a superficial task and it is essential for the effective use of this type of archives. The standardization system that deals with audio-visual descriptors is the MPEG-7 ("Motion Picture Expert Group - 7").

Types of visual descriptors

Descriptors are the first step to find out the connection between pixels contained in a digital image and what humans recall after having observed an image or a group of images after some minutes.

Visual descriptors are divided in two main groups:
# General information descriptors: they contain low level descriptors which give a description about color, shape, regions, textures and motion.
# Specific domain information descriptors: they give information about objects and events in the scene. A concrete example would be face recognition.

General information descriptors

General information descriptors consist of a set of descriptors that covers different basic and elementary features like: color, texture, shape, motion, location and others. This description is automatically generated by means of signal processing.

* COLOR: the most basic quality of visual content. Five tools are defined to describe color. The three first tools represent the color distribution and the last ones describe the color relation between sequences or group of images:
**"Dominant Color Descriptor (DCD)"
**"Scalable Color Descriptor (SCD)"
**"Color Structure Descriptor (CSD)"
**"Color Layout Descriptor (CLD)"
**"Group of frame (GoF)" or "Group-of-pictures (GoP)"

* TEXTURE: also, an important quality in order to describe an image. The texture descriptors characterize image textures or regions. They observe the region homogeneity and the histograms of these region borders. The set of descriptors is formed by:
**"Homogeneous Texture Descriptor (HTD)"
**"Texture Browsing Descriptor (TBD) "
**"Edge Histogram Descriptor (EHD)"

* SHAPE: contains important semantic information due to human’s ability to recognize objects through their shape. However, this information can only be extracted by means of a segmentation similar to the one that the human visual system implements. Nowadays, such a segmentation system is not available yet, however there exists a serial of algorithms which are considered to be a good approximation. These descriptors describe regions, contours and shapes for 2D images and for 3D volumes. The shape descriptors are the following ones:
**"Region-based Shape Descriptor (RSD)"
**"Contour-based Shape Descriptor (CSD)"
**"3-D Shape Descriptor (3-D SD)"

* MOTION: defined by four different descriptors which describe motion in video sequence. Motion is related to the objects motion in the sequence and to the camera motion. This last information is provided by the capture device, whereas the rest is implemented by means of image processing. The descriptor set is the following one:
**"Motion Activity Descriptor (MAD)"
**"Camera Motion Descriptor (CMD)"
**"Motion Trajectory Descriptor (MTD)"
**"Warping and Parametric Motion Descriptor (WMD and PMD)"

* LOCATION: elements location in the image is used to describe elements in the spatial domain. In addition, elements can also be located in the temporal domain:
**"Region Locator Descriptor (RLD)"
**"Spatio Temporal Locator Descriptor (STLD)"

pecific domain information descriptors

These descriptors, which give information about objects and events in the scene, are not easily extractable, even more when the extraction is to be automatically done. Nevertheless they can be manually processed.

As mentioned before, face recognition is a concrete example of an application that tries to automatically obtain this information.

Descriptors applications

Among all applications, the most important ones are:
* Multimedia documents search engines and classifiers.
* Digital library: visual descriptors allow a very detailed and concrete search of any video or image by means of different search parameters. For instance, the search of films where a known actor appears, the search of videos containing the Everest mountain, etc.
* Personalized electronic news service.
* Possibility of an automatic connection to a TV channel broadcasting a soccer match, for example, whenever a player approaches the goal area.
* Control and filtering of concrete audio-visual contents, like violent or pornographic material. Also, authorization for some multimedia contents.

ee also

MPEG-7

DSpace

Feature detection

References

B.S. Manjunath (Editor), Philippe Salembier (Editor), and Thomas Sikora (Editor): "Introduction to MPEG-7: Multimedia Content Description Interface". Wiley & Sons, April 2002 - ISBN 0-471-48678-7

External links

*Multimedia Content Analysis Using both Audio and Video Clues [http://vision.poly.edu:8080/~jhuang/Publication/Content_Analysis_Wang2000SP.pdf]

*Relating Visual and Semantic Image Descriptors [http://www.acemedia.org/aceMedia/files/document/wp7/2004/ewimt04-dcuThom.pdf]

*Fusing MPEG-7 visual descriptors for image classication [http://www.acemedia.org/aceMedia/files/document/wp7/2005/icann05-iti.pdf]

*MPEG-7 Quick Reference [http://gondolin.rutgers.edu/MIC/text/how/mpeg7ref.pdf]


Wikimedia Foundation. 2010.

Игры ⚽ Нужно сделать НИР?

Look at other dictionaries:

  • Wine tasting descriptors — The use of wine tasting descriptors allows the taster an opportunity to put into words the aromas and flavors that they experience and can be used in assessing the overall quality of wine. Many wine writers, like Karen MacNeil in her book The… …   Wikipedia

  • Color Layout Descriptor — (CLD) is designed to capture the spatial distribution of color in an image. The feature extraction process consists of two parts; grid based representative color selection and Discrete Cosine Transform with quantization. Color is the most basic… …   Wikipedia

  • Descriptores visuales — Saltar a navegación, búsqueda Los descriptores visuales describen las características visuales de los contenidos dispuestos en imágenes o en vídeos. Describen características elementales tales como la forma, el color, la textura o el movimiento,… …   Wikipedia Español

  • Automatic image annotation — (also known as automatic image tagging) is the process by which a computer system automatically assigns metadata in the form of captioning or keywords to a digital image. This application of computer vision techniques is used in image retrieval… …   Wikipedia

  • MPEG-7 — consiste en una representación estándar de la información audiovisual que permite la descripción de contenidos (metadatos) para: Palabras clave Significado semántico (quién, qué, cuándo, dónde) Significado estructural (formas, colores, texturas,… …   Wikipedia Español

  • Blob detection — Feature detection Output of a typical corner detection algorithm …   Wikipedia

  • Scale-invariant feature transform — Feature detection Output of a typical corner detection algorithm …   Wikipedia

  • Scale-invariant feature transform — Exemple de résultat de la comparaison de deux images par la méthode SIFT (Fantasia ou Jeu de la poudre, devant la porte d’entrée de la ville de Méquinez, par Eug …   Wikipédia en Français

  • MPEG-7 — is a multimedia content description standard. It was standardized in ISO/IEC 15938 (Multimedia content description interface).[1][2][3][4] This description will be associated with the content itself, to allow fast and efficient searching for… …   Wikipedia

  • information processing — Acquisition, recording, organization, retrieval, display, and dissemination of information. Today the term usually refers to computer based operations. Information processing consists of locating and capturing information, using software to… …   Universalium

Share the article and excerpts

Direct link
Do a right-click on the link above
and select “Copy Link”