Skip to main navigation Skip to search Skip to main content

Visual semantics: extracting visual information from text accompanying pictures

  • SUNY Buffalo

Research output: Contribution to conferencePaperpeer-review

38 Scopus citations

Abstract

This research explores the interaction of textual and photographic information in document understanding. The problem of performing general-purpose vision without a priori knowledge is difficult at best. The use of collateral information in scene understanding has been explored in computer vision systems that use scene context in the task of object identification. The work described here extends this notion by defining visual semantics, a theory of systematically extracting picture-specific information from text accompanying a photograph. Specifically, this paper discusses the multi-stage processing of textual captions with the following objectives: (i) predicting with objects (implicitly or explicitly mentioned in the caption) are present in the picture and (ii) generating constraints useful in locating/identifying these objects. The implementation and use of a lexicon specifically designed for the integration of linguistic and visual information is discussed. Finally, the research described here has been successfully incorporated into PICTION, a caption-based face identification system.

Original languageEnglish
Pages793-798
Number of pages6
StatePublished - 1994
EventProceedings of the 12th National Conference on Artificial Intelligence. Part 1 (of 2) - Seattle, WA, USA
Duration: Jul 31 1994Aug 4 1994

Conference

ConferenceProceedings of the 12th National Conference on Artificial Intelligence. Part 1 (of 2)
CitySeattle, WA, USA
Period07/31/9408/4/94

Fingerprint

Dive into the research topics of 'Visual semantics: extracting visual information from text accompanying pictures'. Together they form a unique fingerprint.

Cite this