GeThR-Net: A Generalized Temporally Hybrid Recurrent Neural Network for Multimodal Information Fusion
Ankit Gandhi, Arjun Sharma, Arijit Biswas, Om Deshmukh
In this paper, we propose a novel generalized deep neural network architecture where temporal streams from multiple modalities can be combined.There are total M+1 (M is the number of modalities) components in the proposed network. The first component is a novel temporally hybrid RNN that exploits the complimentary nature of the multimodal temporal information by allowing the network to learn both modality specific temporal dynamics as well as the dynamics in a multimodal feature space. M additional components are added to the network which extract discriminative but non-temporal cues from each modality. Finally, the predictions from all of these components are linearly combined using a set of automatically learned weights. We perform exhaustive experiments on three different datasets spanning four modalities. The proposed network is relatively 3.5%, 5.7% and 2% better than the best performing temporal multimodal baseline for UCF-101, CCV and Multimodal Gesture datasets respectively. pdf. code
Accepted at ECCV 2016 workshop on Computer Vision for Audio-Visual Media
LIVELINET: A Multimodal Deep Recurrent Neural Network to Predict Liveliness in Educational Videos
Arjun Sharma, Arijit Biswas, Ankit Gandhi, Sonal Patil, Om Deshmukh
Online educational videos have emerged as one of the most popular modes of learning in the recent years. Studies have shown that liveliness is highly correlated to engagement in educational videos. While previous work has focused on feature engineering to estimate liveliness and that too using only the acoustic information, in this paper we propose a technique based on the combination of audio and visual features to predict liveliness. We also propose a novel multimodal deep recurrent neural network based approach to automatically estimate if an educational video is lively or not. On the StyleX dataset of 450 one-minute long educational video snippets, our approach shows an absolute improvement of 8.9% in the classification accuracy over the baseline. pdf
Accepted as Oral at EDM 2016. Nominated for Best Paper Award
Enhancing RGB CNNs with Depth
Arjun Sharma, Pramod Sankar K.
Most current approaches for recognition in RGB-D images fall in either the late fusion or the early fusion category. A drawback of the early fusion scheme is its inapplicability when one of the modalities is absent at test time. On the other hand, a late fusion of features does not allow the correlated nature of modalities to be exploited effectively. Recent approaches using Deep Learning are not immune to these problems either. In this project, we proposed a simple, yet elegant method towards combining early and late fusion of colour and depth information when training Deep Convolutional Neural Networks. We show that when fine-tuning CNNs, an intermediate depth pre-training step provides a significant jump in colour recognition accuracy. pdf ppt
Oral at ACPR 2015, Patent pending
Adapting Off-the-Shelf CNNs for Word Spotting and Recognition
Arjun Sharma, Pramod Sankar K.
The word spotting approach is extremely useful for searching and annotating documents for which robust recognizers are unavailable. Traditionally, hand-designed features were used to represent the word images for spotting. In this paper, we learn a data-driven representation for word-images from Convolutional Neural Networks (CNNs). Previous approaches that learn deep neural networks for a particular task/dataset are difficult to design and train for generic word spotting. Instead, by “adapting” a CNN trained for a different problem, we show tremendous speedup in the training phase. Our experiments show that features extracted from an adapted-CNN handsomely outperform hand-designed features on both spotting and recognition tasks for printed (English and Telugu) and handwritten (IAM) document collections. pdf
Oral at ICDAR 2015
Focus estimation in images
Arjun Sharma, Pramod Sankar K
In this project, we used a pre-trained Deep Convolutional Neural Network to identify pixels in an image which are in focus. We trained a network on a dataset of wildlife images collected from the web and showed superior performance as compared to other recent methods. We then used this network to find focused regions in videos of house inspection recorded using drones and found the performance on this different domain to be surprisingly good.
Patent pending
Emotion recognition in the wild
Guide: Prof. Aaron Courville, Prof. Yoshua Bengio
The goal of the Emotion Recognition in the Wild Challenge 2013 was to develop a system to classify emotions expressed by the primary human subject in the short video clips of acted scenes lasting approximately one-two seconds, including the audio track which may contain human voices as well as background music.
I was a member of Université de Montréal's team that participated in this challenge. I worked on adapting Wang et al's activity recognition using dense trajectories pipeline for emotion recognition. I was also involved in generating a video-emotion label dataset from descriptive video service files of movies. pdf
Accepted at ICMI 2013
Evaluation of Multi-class SVMs using DEA
Guide: Prof. V. Vijaya Saradhi
One-vs-One, One-vs-All, and Error correcting output codes (ECOC) are some of the most widely employed methods for multi-class classification using binary classifiers. The superiority of any one method over the others has been the subject of much research. In this project, we compared these methods on the basis of parameters that describe an optimal separating hyperplane in the best possible way, namely the training error, the S-span, the fraction of support vectors and the margin of individual hyperplanes unlike the traditional method of comparing them along the test error. We proposed a novel method of using Data Envelopment Analysis (DEA) to compute the relative efficiencies of the hyperplanes instead of comparing the entire ensemble as one. pdf
Developing a Machine Translation System using Translation Memory
Guide: Dr. Sandipan Dandapat
The aim of the project was to improve upon present Translation memory systems. Given an input sentence, the search for the closest match in the translation memory in the source language was improved using an IR engine and tree based methods as compared to the traditional edit distance based approaches. The difference in the source sentence, its closest match and the translation in the memory was identified and automatic post-editing carried out using statistical machine translation to obtain an accurate translation for the source sentence. ppt1 ppt2
Activity Recognition in Egocentric Videos using Deep Learning
Vardaan Pahuja, Arjun Sharma, Ankit Gandhi, Arijit Biswas
In this project, we compared the performance of state-of-the-art deep learning methods against previous techniques on the activities of daily living (ADL) dataset of first person videos. We compared R-CNN object detectors with previously used DPM based object detectors. We also compared the performance of Long-Short Memory (LSTM) networks with the previously used temporal pyramid based approach to aggregate per frame object detection scores for video classification.