Abstract Facial emotion recognition (FER) is one of the main fields of research in image processing and artificial intelligence with applications in human–machine interaction, behavior analysis, and intelligent vision. The RAF DB dataset, a widely used benchmark dataset for emotion recognition, is used for the training and testing of deep learning models. In real application conditions, the facial image may be subjected to visual noise. Large glasses that occlude the eye region of the face are one example of environmental or artificial noise which distorts the visual information. If frames covering a large area of an expressive part of the face, like the eyes, are recognized as noise by the emotion recognition network this may lead to noticeable performance degradation. Artificial occlusion noise, similar to large glasses, was applied to the eye region of the facial images from the RAF DB dataset. The impact of this type of structural noise on the performance of deep-learning-based facial emotion recognition models is analyzed. Three state-of-the-art architectures (ResNet-50, Vision Transformer (ViT-Base/16) and Swin Transformer Tiny (Swin T)) are tested to determine their robustness to the noisy application conditions. The objective of the study is to determine and suggest possible methods to counter the decrease in recognition performance caused by the occlusion noise. Experimental results indicate that, within our evaluation setup, transformer-based architectures, especially Swin-T, tend to be more robust to the simulated visual noise, which may be useful for the design of facial emotion recognition models intended for more complex conditions. Similar content being viewed by others Funding No funding was received for this work. Author information Authors and Affiliations Corresponding author Ethics declarations Competing interests The authors declare no competing interests. Ethical and informed consent This article does not contain any studies with human participants or animals performed by any of the