Video captioning of object related activities using estimation of object shift vectors

Authors

  • Prashant Kaushik ,Vikas Saxena,Amarjeet Prajapati

Keywords:

Geometric transformations, monitoring several objects, periodic motion, video captioning, estimating speed, and video captioning, objected oriented video captioning, video understanding.

Abstract

Object-oriented video captioning focuses on generating descriptions, summarization that emphasize object-related activities, including basic interactions such as object-object dynamics their interection with background and engagement with the background. Traditional methods primarily provide general video descriptions by extracting spatial and temporal features but often fail to capture how objects interact. To address this gap, our study examines object movements and trajectories to estimate interactions within a scene.

References

Lin, K., Li, L., Lin, C.C., Ahmed, F., Gan, Z., Liu, Z., Lu, Y. and Wang, L., 2022. Swinbert: End-to-end transformers with sparse attention for video captioning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 17949-17958).

P. Kaushik, K. V. Kumar and P. Biswas, "Context Bucketed Text Responses using Generative Adversarial Neural Network in Android Application with Tens or Flow-Lite Framework," 2022 8th International Conference on Signal Processing and Communication (ICSC), Noida, India, 2022, pp. 324-328, doi: 10.1109/ICSC56524.2022.10009634.

Downloads

Published

2024-12-10

How to Cite

Prashant Kaushik ,Vikas Saxena,Amarjeet Prajapati. (2024). Video captioning of object related activities using estimation of object shift vectors . Journal of Computational Analysis and Applications (JoCAAA), 33(08), 2723–2738. Retrieved from https://eudoxuspress.com/index.php/pub/article/view/2197

Issue

Section

Articles