Vision–language models · Multilingual NLP
IMMCAN
Connecting idioms, captions, and images
Ranking five image–caption candidates by how well they express the intended meaning of an idiom in context.
My contribution
Equal-contribution co-author with Barış Bilen · Multimodal framework development
I co-developed a multimodal cross-attention framework integrating contextual language representations with visual and caption features for idiom understanding.

01 / Problem
What needed solving
Literal image matching can miss an idiom’s intended meaning. The task also requires transfer across languages with limited supervision.
02 / Approach
Methods & data
AdMIRe 1.0 for supervised development; AdMIRe 2.0 for multilingual zero-shot evaluation. MAGPIE supports idiomaticity detection.
- Encode the contextual idiom and image–caption candidates
- Fuse caption and image features with cross-attention
- Condition on the idiom and rank candidates
03 / Outcome
35.0% zero-shot top-image accuracy
On the AdMIRe 2.0 ALL row, VTT-Cla-Base reached 0.350 accuracy and 0.727 NDCG, versus 0.288 and 0.704 for TT-Cla-Base (Table 3).
The text-only model performed better on the smaller AdMIRe 1.0 test. Caption augmentation had mixed effects. The linked repository currently presents the project overview; implementation availability should be checked there.
Read the published evaluationPublication
VisAffect at MWE-2026 AdMIRe 2: IMMCAN Idiom Multimodal Cross-Attention Network
Proceedings of the 22nd Workshop on Multiword Expressions (MWE 2026) · 2026