Ali Azmoudeh /

Vision–language models · Multilingual NLP

IMMCAN

Connecting idioms, captions, and images

Ranking five image–caption candidates by how well they express the intended meaning of an idiom in context.

My contribution

Equal-contribution co-author with Barış Bilen · Multimodal framework development

I co-developed a multimodal cross-attention framework integrating contextual language representations with visual and caption features for idiom understanding.

Conceptual connections between visual and language representations
Conceptual illustration · not experimental output

01 / Problem

What needed solving

Literal image matching can miss an idiom’s intended meaning. The task also requires transfer across languages with limited supervision.

02 / Approach

Methods & data

  • XLM-R
  • Jina-CLIP-v2
  • Frozen pretrained encoders
  • Two-stage cross-attention
  • Ranking objectives
  • Caption augmentation

AdMIRe 1.0 for supervised development; AdMIRe 2.0 for multilingual zero-shot evaluation. MAGPIE supports idiomaticity detection.

  1. Encode the contextual idiom and image–caption candidates
  2. Fuse caption and image features with cross-attention
  3. Condition on the idiom and rank candidates

03 / Outcome

35.0% zero-shot top-image accuracy

On the AdMIRe 2.0 ALL row, VTT-Cla-Base reached 0.350 accuracy and 0.727 NDCG, versus 0.288 and 0.704 for TT-Cla-Base (Table 3).

The text-only model performed better on the smaller AdMIRe 1.0 test. Caption augmentation had mixed effects. The linked repository currently presents the project overview; implementation availability should be checked there.

Read the published evaluation

Publication

VisAffect at MWE-2026 AdMIRe 2: IMMCAN Idiom Multimodal Cross-Attention Network

Barış Bilen, Ali Azmoudeh, Hazım Kemal Ekenel, Hatice Köse

Proceedings of the 22nd Workshop on Multiword Expressions (MWE 2026) · 2026