PhD Student and Researcher · Computer Vision

Ali Azmoudeh

I develop computer vision systems that remain reliable when data is limited, synthetic, or shifted—across facial analysis, medical imaging, and multimodal visual understanding.

Multimodal vision / 02Visual + Language
Computer vision graphic combining facial landmark analysis with brain imaging analysis
FACEvisual analysis
BRAINmedical imaging
VISION–LANGUAGEmultimodal reasoning
01 / PROFILE

Engineering vision
for the real world.

I am a PhD student in Computer Engineering at Istanbul Technical University and a Graduate Research Assistant at SiMiT Lab, working under the supervision of Prof. Dr. Hazım Kemal Ekenel.

My research sits at the intersection of computer vision, deep learning, and trustworthy AI. I study how vision models behave beyond clean benchmarks: under domain shift, limited annotations, low-quality medical scans, privacy constraints, and imbalanced data.

Current directions include reliable brain-tumor MRI segmentation, conformal risk control, vision-language-guided facial expression generation, multimodal idiom understanding, and class-agnostic visual counting.

My technical toolkit also includes large language models (LLMs), Hugging Face Transformers, and prompt engineering, alongside AI-assisted development with Claude Code, OpenAI Codex, ChatGPT, and Gemini.

Based
Istanbul, Türkiye
Affiliation
ITU · SiMiT Lab
Role
PhD Researcher · RA/TA
Advisor
Prof. Dr. Hazım Kemal Ekenel
Focus
Computer Vision · Multimodal AI · Trustworthy Medical AI
02 / RESEARCH

Three connected
research directions.

From pixels to decisions, my work asks the same question: how can a model generalize responsibly when the real world does not match its training set?

[ MED / 3D ]

Medical image analysis

Brain-tumor segmentation across high- and low-quality MRI domains, combining strong 3D architectures with transfer learning and uncertainty-aware evaluation.

[ FACE / SYN ]

Facial analysis & synthetic data

Privacy-aware facial expression recognition and text-guided expression generation using vision-language and generative models, synthetic data, and cross-dataset evaluation.

[ VISION / LANGUAGE ]

Multimodal visual understanding

Vision-language models (VLMs) and visual-textual cross-attention for multimodal reasoning, including IMMCAN idiom understanding and open-world text-guided visual counting.

03 / OUTPUT

Selected publications.

Peer-reviewed work and open research across medical vision, facial expression recognition, multimodal reasoning, and visual counting.

P / 04

VisAffect at MWE-2026 AdMIRe 2: IMMCAN Idiom Multimodal Cross-Attention Network

22nd Workshop on Multiword Expressions · ACL, 2026 · pp. 149–153

Integrates visual and textual representations through multimodal cross-attention for idiom understanding and multimodal reasoning.

P / 03

On Applicability of Synthetic Datasets for Facial Expression Recognition

IEEE FG 2026 · Kyoto, Japan · First author

Examines pseudo-labeling, diffusion-based synthesis, and GAN-based expression editing as privacy-preserving ways to address imbalance and limited facial-expression data.

P / 02

GLIMS-MedNeXt: An Ensemble Framework for Brain MRI Segmentation in Sub-Saharan Africa

MICCAI 2025 BraTS-Lighthouse · LNCS 16376 · Springer, 2026

Combines GLIMS and MedNeXt with transfer learning and ensemble fusion to improve brain-tumor segmentation under low-quality imaging conditions.

P / 01

A Survey on Class-Agnostic Counting: Advancements from Reference-Based to Open-World Text-Guided Approaches

Computer Vision and Image Understanding · 2026 · Article 104703

Introduces a taxonomy spanning reference-based, reference-less, and open-world text-guided counting, and reviews 29 approaches on established benchmarks.

04 / PATH

Research & education.

2025—

PhD, Computer Engineering

Istanbul Technical University · Computer vision and trustworthy medical AI

2024—

Turkcell Graduate Research & Teaching Assistant

Istanbul Technical University · Research and undergraduate teaching

2022—

Graduate Research Assistant

SiMiT Lab, ITU · Computer vision, biometrics, and medical imaging

2022–25

MSc, Computer Engineering

Istanbul Technical University

2016–21

BSc, Computer Engineering

University of Tabriz

Research collaboration

Let’s build vision systems that generalize.

For research discussions, collaborations, or questions about my published work, reach me through my ITU email.

Start a conversation