Video captioning of object related activities using estimation of object shift vectors
Keywords:
Geometric transformations, monitoring several objects, periodic motion, video captioning, estimating speed, and video captioning, objected oriented video captioning, video understanding.Abstract
Object-oriented video captioning focuses on generating descriptions, summarization that emphasize object-related activities, including basic interactions such as object-object dynamics their interection with background and engagement with the background. Traditional methods primarily provide general video descriptions by extracting spatial and temporal features but often fail to capture how objects interact. To address this gap, our study examines object movements and trajectories to estimate interactions within a scene.
References
Lin, K., Li, L., Lin, C.C., Ahmed, F., Gan, Z., Liu, Z., Lu, Y. and Wang, L., 2022. Swinbert: End-to-end transformers with sparse attention for video captioning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 17949-17958).
P. Kaushik, K. V. Kumar and P. Biswas, "Context Bucketed Text Responses using Generative Adversarial Neural Network in Android Application with Tens or Flow-Lite Framework," 2022 8th International Conference on Signal Processing and Communication (ICSC), Noida, India, 2022, pp. 324-328, doi: 10.1109/ICSC56524.2022.10009634.


