11-682 Visual Learning and Recognition
Multi-modal systems are the closest approximation to human perception, hence working on multimodal systems that work with conversational phrases is essential for the improvement of the interaction between humans and machines and the integration of AI into everyday life. Our idea is related to the grounding of objects and actions in images. This basically involves taking a prompt from a user and finding the objects/actions corresponding to the prompt from the visual input.
This can have a range of applications, including use in the second generation of personal assistants (Siri, Alexa, etc.) which can operate in a multimodal setting (audio+visual input). Another interesting application for this field of work (when extended to videos) is the automatic annotation of a video (Youtube video), either from the content creators end or the content consumers end.
Integration of AI in everyday life through humans and machine interaction
Next-generation personal assistants like Alexa, Siri, etc.