Abstract
Pre-training visual and textual representations from large-scale image-text pairs is becoming a standard approach for many downstream vision-language tasks. The transformer-based models learn inter- and intramodal attention through a list of self-supervised learning tasks. This paper proposes LAViTeR, a novel architecture for visual and textual representation learning. The main module, Visual Textual Alignment (VTA) will be assisted by two auxiliary tasks, GAN-based image synthesis and Image Captioning. We also propose a new evaluation metric measuring the similarity between the learnt visual and textual embedding. The experimental results on two public datasets, CUB and MS-COCO, demonstrate superior visual and textual representation alignment in the joint feature embedding space. Our code is publicly available at https://github.com/mshaikh2/MMRL.
| Original language | English |
|---|---|
| Pages (from-to) | 162-169 |
| Number of pages | 8 |
| Journal | IET Conference Proceedings |
| Volume | 2024 |
| Issue number | 10 |
| DOIs | |
| State | Published - 2024 |
| Event | 26th Irish Machine Vision and Image Processing Conference, IMVIP 2024 - Limerick, Ireland Duration: Aug 21 2024 → Aug 23 2024 |
Fingerprint
Dive into the research topics of 'LAViTeR: Learning Aligned Visual and Textual Representations Assisted by Image and Caption Generation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver