Efthymia Tsamoura
Eight Years of Weakly Supervised Neurosymbolic Learning: Main Results and Algorithms
Neurosymbolic learning (NSL)—the integration of neural and symbolic mechanisms for inference and learning—has been proposed as a remedy for some of the most critical limitations of neural networks. A problem that has received considerable attention in the relevant literature is that of weakly supervised learning in the presence of a symbolic component. In this problem, one or more neural classifiers perform symbolic grounding, mapping low-level inputs onto high-level symbolic concepts. A symbolic component then infers outputs consistent with both the predicted concepts and any provided prior knowledge, and learning is driven by supervision applied only to the outputs of the symbolic component. Research has uncovered an intriguing pitfall. In particular, symbol grounding can "deceive" the learning process: even when the neural classifiers produce incorrectly grounded concepts, the symbolic component may still infer outputs that match the ground truth.
This talk provides an overview of research on weakly supervised NSL over the last eight years, with a focus on the above challenge. We begin by discussing early learning frameworks, then present the main theoretical results—including PAC-learnability, reasoning shortcuts, and learning imbalances. We conclude with an overview of strategies for improving learning efficiency, delving into a specific technique that enhances the accuracy of learned classifiers by up to 53% through better exploiting the representation space.