Friday, August 21 : 8h30-12h30
8h30 - 8h45 Opening remarks
8h45 - 9h10 Summarization of egocentric video using graphs by Pr. Abhimanyu Sahu
In this session, we show how different graph representations can be developed for accurately summarizing first-person (egocentric) videos in a computationally efficient manner. Each frame in a video is first represented as a weighted graph. We develop a new shot boundary detection method using graph based mutual information models. We next construct a weighted graph for each shot. A representative frame from each shot is selected using a graph centrality measure. A new way of characterizing egocentric video frames using a graph-based center- surround model is shown next. Here, each representative frame is modeled as a union of two graphs: one for the center region and the other for the surround region. By exploiting spectral measures of dissimilarity between the two (center and surround) graphs, optimal center and surround regions are determined. Optimal regions for all frames within a shot are kept the same as that of the representative frame. Center-surround differences in entropy and optical flow values along with PHOG (Pyramidal HOG) features are extracted from each frame. All frames in a video are finally represented by another weighted graph, termed as a Video Similarity Graph (VSG). The frames are clustered by applying a Minimum Spanning Tree (MST) based approach with a new measure for inadmissible edges. Frames closest to the centroid of each cluster are captured to build the summary. Comprehensive comparisons with several state-of-the-art methods clearly indicate the benefit of our framework.
9h10- 9h25 Practices I
9h25- 9h50 Action recognition and localization in egocentric video under unimodal and multimodal settings with random walks on graphs and deep learning by Pr. Abhimanyu Sahu
In this session, we propose two solutions to the problem action recognition in ego centric videos. As the first solution, we apply random walks to recognize a single action. In a more comprehensive second solution, we propose a superpixel level weakly supervised graph-theoretic framework for joint localization, recognition and summarization of actions in an egocentric video. Here, we first recognize and localize single as well as multiple action(s) in each frame of an egocentric video and then construct a summary of these detected actions. The superpixel level solution helps in precise localization of actions in addition to improving the recognition accuracy. After determining action label(s) for each frame from its constituent superpixels, we apply a fractional knapsack type formulation for obtaining a summary (of actions). Experimental comparisons on 5 publicly available datasets show the effectiveness of the proposed solution. Our solutions require only a few training data (seeds) as compared to existing solutions which mostly require a considerably large number of training data.
9h50- 10h10 Practices II
10h10- 10h30 On Video Analysis from Event Cameras with graph based methods by Pr. Ananda S. Chowdhury
10h30- 11h00 Coffee Break
11h00- 11h20 Fundamentals of graph-based models (graph signal processing, graph neural networks and hypergraphs) by Ass. Prof. Anastasia Zakharova
11h20- 11h40 Video data analysis with GNNs (transductive and inductive models) for surface environments with conventional cameras by Ass. Prof. Anastasia Zakharova
11h40- 12h00 Video data analysis with GNNS for underwater environments with conventional cameras by Dr. Meghna Kapoor
Underwater moving object detection is important for monitoring marine habitats and supporting the study of marine organisms, yet it remains challenging because of dynamic backgrounds, uneven illumination, and changes in object appearance. In this session, we discuss how graph-based representations can be used to analyze underwater video and improve the segmentation of moving objects from complex backgrounds. By capturing relationships among visual, motion, and appearance information, graph learning offers a useful framework for handling the variability commonly observed in marine scenes. The session highlights the role of such approaches in fish detection and segmentation and their broader potential for passive, camera-based monitoring of underwater environments.
12h00- 12h25 Practices III
12h25- 12h30 Concluding remarks