Prof. Bahareh Tolooshams, University of Alberta
YouTube Stream: https://www.youtube.com/watch?v=7yZvJCN_qyM
Join group to receive calendar invite: https://groups.google.com/a/modelingtalks.org/g/talks
Abstract:
Modern AI systems are shaped by their inductive biases, which implicitly define the structure of the representations they learn. These biases become particularly important when models are used for scientific understanding, where the goal is not only prediction, but extracting meaningful and potentially causal structure from data. In this talk, I present a perspective in which learning is viewed as inference, linking representation learning, generative modelling, computational neuroscience, and inverse problems. I revisit the popular linear representation hypothesis used in the mechanistic interpretability literature and discuss the assumptions it imposes about the geometry of internal representations, rather than discovering their underlying geometric structure. I then discuss extracting hierarchical and functional representations, illustrating how the inference perspective allows us to explore different inductive biases that reveal distinct geometric structures. These results highlight the need for principled, theoretically grounded interpretability methods that better align with the structure of data and learned representations.
Bio:
Dr. Tolooshams is an Assistant Professor at the University of Alberta, a fellow of Alberta Machine Intelligence Institute (Amii), an affiliate of the Neuroscience and Mental Health Institute (NMHI), and a Canada CIFAR AI Chair. Dr. Tolooshams received her PhD in 2023 from Harvard University, where she was also an affiliate to the Center for Brain Science. Before joining the University of Alberta, Dr. Tolooshams was a postdoctoral researcher and held the Swartz Foundation Fellowship in Theoretical Neuroscience for two years at Caltech. Dr. Tolooshams’ research broadly focuses on representation learning, generative models, interpretability, and NeuroAI. She studies how AI systems learn internal representations of the world and how these representations can be structured, disentangled, and interpretable. A central theme of her research is mechanistic interpretability: developing methods to understand how concepts are represented inside modern neural networks and the principles that shape these representations.