Skip to main navigation Skip to search Skip to main content

LAViTeR: Learning Aligned Visual and Textual Representations Assisted by Image and Caption Generation

  • Mohammad Abuzar Hashemi
  • , Zhanghexuan Li
  • , Mihir Chauhan
  • , Yan Shen
  • , Abhishek Satbhai
  • , Mir Basheer Ali
  • , Mingchen Gao
  • , Sargur Srihari
  • SUNY Buffalo

Research output: Contribution to journalConference articlepeer-review

Abstract

Pre-training visual and textual representations from large-scale image-text pairs is becoming a standard approach for many downstream vision-language tasks. The transformer-based models learn inter- and intramodal attention through a list of self-supervised learning tasks. This paper proposes LAViTeR, a novel architecture for visual and textual representation learning. The main module, Visual Textual Alignment (VTA) will be assisted by two auxiliary tasks, GAN-based image synthesis and Image Captioning. We also propose a new evaluation metric measuring the similarity between the learnt visual and textual embedding. The experimental results on two public datasets, CUB and MS-COCO, demonstrate superior visual and textual representation alignment in the joint feature embedding space. Our code is publicly available at https://github.com/mshaikh2/MMRL.

Original languageEnglish
Pages (from-to)162-169
Number of pages8
JournalIET Conference Proceedings
Volume2024
Issue number10
DOIs
StatePublished - 2024
Event26th Irish Machine Vision and Image Processing Conference, IMVIP 2024 - Limerick, Ireland
Duration: Aug 21 2024Aug 23 2024

Fingerprint

Dive into the research topics of 'LAViTeR: Learning Aligned Visual and Textual Representations Assisted by Image and Caption Generation'. Together they form a unique fingerprint.

Cite this