Integrated Cross-Modal Understanding through Multimodal Large Language Models

Authors

  • Chintu Kodanda Ramu , Dr.Pankaj Khairnar

Keywords:

Multimodal Large Language Models, Multimodal Learning, Cross-Modal Representation Learning, Transformer Architectures, Text-Image-Speech Integration, Attention Mechanisms, Deep Learning, Artificial Intelligence, Multimodal Fusion, Context-Aware AI

Abstract

The rapid advancement of Artificial Intelligence (AI) has significantly transformed the abilityof machines to process and interpret complex information [1]. Traditional Large LanguageModels (LLMs) have achieved remarkable success in natural language understanding and
generation; however, their capabilities remain largely restricted to textual data [2]. Human cognition, in contrast, operates

References

Vaswani, A., et al., “Attention Is All You Need,” Advances in Neural Information Processing Systems, 2017.

Devlin, J., et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” NAACL, 2019

Downloads

Published

2024-10-15

How to Cite

Chintu Kodanda Ramu , Dr.Pankaj Khairnar. (2024). Integrated Cross-Modal Understanding through Multimodal Large Language Models . Journal of Computational Analysis and Applications (JoCAAA), 33(08), 8697–8705. Retrieved from https://eudoxuspress.com/index.php/pub/article/view/5432

Issue

Section

Articles