September 9, 2026
Malmö, Sweden
September 9, 2026
Malmö, Sweden
Real-world deployment introduces challenges that benchmarks often overlook, including limited compute, strict latency requirements, domain shifts, and noisy or incomplete data. This workshop addresses these gaps by focusing on Real-World Video Representation Learning: transitioning from evaluation-driven development to model representations designed for real-world deployment.
By shifting the focus from leaderboard-driven improvements to reliability, efficiency, adaptability, and real-world generalization, the workshop aims to bridge cutting-edge research in video representation learning with the operational demands of real-world AI systems. Overall, it promotes a research agenda that prioritizes impact, deployability, and sustained performance beyond controlled benchmarks.
INVITED SPEAKERS & PANELISTS
Google DeepMind, UK
Ecole des Ponts ParisTech, FR
AGENDA
09:00 - 09:10: Introduction
09:10 - 09:40: Speaker 1: Hazel Doughty
09:40 - 10:10: Speaker 2: Viorica Patraucean
10:10 - 10:30: Coffee Break (Malmo Massan Exhibit Hall)
10:30 - 11:00: Oral Session
11:00 - 11:30: Speaker 3: Gül Varol
11:30 - 12:00: Panel Discussion [Hazel Doughty, Yoichi Sato, Gül Varol, Ye Xia]
12:00 - 12:05: Closing Remarks
12:05 - 12:30: Break
12:30 - 13:30: Poster Session
Are We Making Progress in Multimodal Domain Generalization? A Comprehensive Benchmark Study, Hao Dong
CoPE-VideoLM: Leveraging Codec Primitives For Efficient Video Language Modeling, Sayan Deb Sarkar, Rémi Pautrat, Ondrej Miksik, Marc Pollefeys, Iro Armeni, Mahdi Rad, Mihai Dusmanu [Oral]
Modality-Aware Out-of-Distribution Detection for Multi-Modal Action Recognition, Lars Doorenbos, Duc Manh Vu, Serday Ozsoy, Juergen Gall
Toward Real-World Neural Video Representation: Compact, Low-Complexity, and Scalable, Ho Man Kwan, Tianhao Peng, Fan Zhang, David Bull
VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement, Seohyun Lee, Seoung Choi, Dohwan Ko, Jongha Kim, Hyunwoo Kim [Oral]
TTA-Vid: Generalized Test-Time Adaptation for Video Reasoning, Soumya Jahagirdar, Edson Araujo, Anna Kukleva, Jehanzeb Mirza, Saurabhchand Bhati, Samuel Thomas, Brian Kingsbury, Rogerio Feris, James Glass, Hilde Kuehne [Oral]
What Moves? Context-Aware Localized Latent Actions, Frank Fundel, Malek Ben Alaya, Thomas Ressler-Antal, Stefan Andreas Baumann, Björn Ommer
When Conditional Sequence Matching Does Not Transfer to Global Video Retrieval, Arjang Talattof
VidEoMT: Your ViT is Secretly Also a Video Segmentation Model, Narges Norouzi, Idil Esen Zulfikar, Niccolò Cavagnero, Tommie Kerssies, Bastian Leibe, Gijs Dubbelman, Daan de Geus
I Have a Stream: Making Self-Supervised Learning Work on Continuous Video, Ivan Martinović, Lukas Knobel, Yuki M. Asano
Adapting MLLMs for Nuanced Video Retrieval, Piyush Bagad, Andrew Zisserman
Please prepare the posters following the official ECCV 2026 poster guidelines.
ORGANIZING TEAM
University of Twente, NL
University of Twente, NL
Eindhoven University of Technology, NL
Bocconi University, IT
PROGRAM COMMITTEE
Ana Manzano - UvA / Amsterdam UMC
Melissa Tijink - University of Twente
Mohammadreza Salehi - University of Amsterdam
Niccolò Cavagnero - Eindhoven University of Technology
Rakshith Srinivasa Murthy - Adobe
Riccardo Santambrogio - Politecnico di Milano
Simone Alberto Peirone - Politecnico di Torino
Zhiqi Miao - University of Groningen