Evaluating affect analysis methods presents challenges due to inconsistencies in database partitioning and evaluation protocols, leading to unfair and biased results. Previous studies claim continuous performance improvements, but our findings challenge such assertions. Using these insights, we propose a unified protocol for database partitioning that ensures fairness and comparability.
For 6 widely used affective databases, we provide detailed demographic annotations (in terms of race, gender and age), evaluation metrics, and a common framework for expression recognition, action unit detection and valence-arousal estimation.
We also provide multiple trained models (baseline and state-of-the-art) with the new protocol and introduce a new leaderboards to encourage future research in affect recognition with a fairer comparison. Our annotations, code, and pre-trained models are available here.
FairFuse a novel debiasing method that yields a closed form solution in the cross-modal space, achieving Pareto optimal fairness with bounded utility losses. This method is training-free, requires no extra data and sensitive attribute annotations, and can debias both visual and textual modalities across multiple tasks The work entitled 'A Closed-Form Solution for Debiasing Vision-Language Models with Utility Guarantees Across Modalities and Tasks' has been published in IEEE CVPR 2026. Code of this development is available here.
FuseMamba-VD a novel method designed for automated violence detection in surveillance and real-world videos. The model uses two VideoMamba branches with different scan orders and performs continuous Gated Class Token Fusion from the spatial-first branch into the temporal-first branch. The work entitled 'FuseMamba-VD: Dual Branch VideoMamba with Gated Class Token Fusion for Violence Detection' has been published in IEEE AVSS 2026. Code of this development is available here.
CUE-Net a novel architecture designed for automated violence detection in video surveillance. CUE-Net combines spatial Cropping with an enhanced version of the UniformerV2 architecture integrating convolutional and self-attention mechanisms alongside a novel Modified Efficient Additive Attention mechanism (which reduces the quadratic time complexity of self-attention) to effectively and efficiently identify violent activities. The work entitled 'CUE-Net: Violence Detection Video Analytics with Spatial Cropping, Enhanced UniformerV2 and Modified Efficient Additive Attention' has been published in IEEE CVPR-W 2024. Code of this development is available here.
A deep neural network for valence-arousal estimation trained on Aff-Wild database, producing state-of-the-art results on it. The trained models can be found here.
This repository contains our solution to the OMG-Emotion Challenge 2018 -using only visual data for valence-arousal estimation utilizing the OMG Emotion database- that ranked 2nd for vision-only valence estimation and 2nd for overall valence estimation. more details with training and inference code can be found here.