Hierarchical Attention Transformers for Multimodal Human Emotion Understanding
Keywords:
Multimodal Emotion Recognition, Transformer Networks, Deep Learning, Cross-Modal Attention, Text-Image-Speech Integration, Affective Computing, Artificial Intelligence, Vision Transformers, Speech Emotion Recognition, Context-Aware LearningAbstract
Human emotions constitute a fundamental component of communication, cognition, andsocial interaction [7]. The increasing demand for emotionally intelligent artificial intelligencesystems has accelerated research in multimodal emotion recognition, where textual, visual, and auditory signals are jointly analyzed to infer emotional states
References
Vaswani, A., et al., “Attention Is All You Need,” NeurIPS, 2017.
Devlin, J., et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” NAACL, 2019.


