This is an advanced course that will focus on the latest research in vision transformers. It consists of research paper presentations, paper discussions, and a semester-long course project. Topics will include transformer architectures for classification and dense prediction tasks in images and videos, efficient transformer architectures, data-efficient learning, self-supervised learning, multi-modal learning, image/video generation, and others. A background in deep learning is required.
Instructor: Gedas Bertasius
Time: Tue & Thu 11:00 am - 12:15 pm
Location: SN 011
Office Hours: By Appointment
TA: Baiqi Li
TA Office Hours: TBD
Canvas Site: link
Class Participation: 10%
Paper Probes: 20%
Paper Presentations: 30%
Course Project: 40%
Class Participation: Please come to class prepared for a paper discussion with your peers.
Late Submissions: The class is structured around a tight paper presentation schedule. Therefore, late assignments will not be accepted.
Academic Integrity: For your presentations and projects, you are allowed to use materials from external sources. However, you must clearly acknowledge those sources.
AI Tool Usage: You are encouraged to use AI tools anywhere in this course, including paper probes, presentations, and projects. Every submission must include a short note listing which tools you used and what you asked them to do. Undisclosed AI use is treated the same as an unacknowledged external source.