Synthetic data · Generative AI
SyntFER
Learning facial expressions from synthetic data
Building and comparing data-generation strategies for facial expression recognition when labels, class balance, and image-sharing constraints matter.
My contribution
First author · Synthetic-data research and cross-dataset evaluation
I investigated three dataset-construction strategies and evaluated their use in facial expression recognition: confidence-based pseudo-labeling, diffusion generation, and GAN expression editing.

01 / Problem
What needed solving
Expression datasets are imbalanced, and collecting or sharing face images raises practical privacy constraints. Can synthetic data improve transfer beyond the training dataset?
02 / Approach
Methods & data
RAF-DB, FER2013, and AffectNet for evaluation; DigiFace, DCFace, EmoNet-Face BIG, and FFHQ as data sources.
- Construct and balance expression datasets
- Train with synthetic-only or mixed data
- Measure cross-dataset accuracy and F1
03 / Outcome
57.02% accuracy on AffectNet
IR50 with Mixed-SYN-C achieved 57.02% accuracy and 56.36% F1 on AffectNet, compared with 40.20% accuracy and 35.82% F1 for the RAF-DB-only baseline (Tables II and IV).
Results depend on the training mixture and target dataset. Synthetic-only training retains a domain gap; these scores are not a universal performance claim.
Read the published evaluationPublication
On Applicability of Synthetic Datasets for Facial Expression Recognition
2026 IEEE 20th International Conference on Automatic Face and Gesture Recognition (FG) · 2026