TY - GEN
T1 - Self-supervised Deformation Modeling for Facial Expression Editing
AU - Athar, Shah Rukh
AU - Shu, Zhixin
AU - Samaras, Dimitris
N1 - Publisher Copyright: © 2020 IEEE.
PY - 2020/11
Y1 - 2020/11
N2 - Deep generative models have recently demonstrated impressive results in photo-realistic facial image synthesis and editing. Existing neural network-based approaches usually only rely on texture generation to edit expressions and largely neglect the motion information. However, facial expressions are inherently the result of muscle movement. In this work, we propose a novel end-to-end network that disentangles the task of facial editing into two steps: a 'motionediting' step and a 'texture-editing' step. In the 'motionediting' step, we explicitly model facial movement through an image deformation, warping the image into the desired expression. In the 'texture-editing' step, we generate the necessary textures, such as teeth and shading effects, for a photorealistic result. Our physically-based task-disentanglement system design allows each step to learn a focused task, and thus need not generate texture to hallucinate motion. Our system is trained in a self-supervised manner, requiring no ground truth deformation annotation. Using Action Units [8] as the representation for facial expression, our method improves the state-of-the-art facial expression editing performance in both qualitative and quantitative evaluations.1.
AB - Deep generative models have recently demonstrated impressive results in photo-realistic facial image synthesis and editing. Existing neural network-based approaches usually only rely on texture generation to edit expressions and largely neglect the motion information. However, facial expressions are inherently the result of muscle movement. In this work, we propose a novel end-to-end network that disentangles the task of facial editing into two steps: a 'motionediting' step and a 'texture-editing' step. In the 'motionediting' step, we explicitly model facial movement through an image deformation, warping the image into the desired expression. In the 'texture-editing' step, we generate the necessary textures, such as teeth and shading effects, for a photorealistic result. Our physically-based task-disentanglement system design allows each step to learn a focused task, and thus need not generate texture to hallucinate motion. Our system is trained in a self-supervised manner, requiring no ground truth deformation annotation. Using Action Units [8] as the representation for facial expression, our method improves the state-of-the-art facial expression editing performance in both qualitative and quantitative evaluations.1.
KW - Expression Editing
KW - Generative Modelling
KW - Self Supervised Learning
UR - https://www.scopus.com/pages/publications/85101452099
U2 - 10.1109/FG47880.2020.00115
DO - 10.1109/FG47880.2020.00115
M3 - Conference contribution
T3 - Proceedings - 2020 15th IEEE International Conference on Automatic Face and Gesture Recognition, FG 2020
SP - 294
EP - 301
BT - Proceedings - 2020 15th IEEE International Conference on Automatic Face and Gesture Recognition, FG 2020
A2 - Struc, Vitomir
A2 - Gomez-Fernandez, Francisco
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 15th IEEE International Conference on Automatic Face and Gesture Recognition, FG 2020
Y2 - 16 November 2020 through 20 November 2020
ER -