Day 1
9:00 - 9:15
9:15 - 10:30
Abstract: In this talk we will discuss the role of causality and abstraction in modelling. We will compare standard modelling in machine learning with modelling relying on the structures of causality and abstraction. After briefly reviewing the formalism of causality and abstraction, we will introduce the problem of causal abstraction learning and we will review selected methods for learning from the literature. We will conclude with a discussion of open questions and relevant future directions of work in this area.
10:30 - 10:45
10:45 - 12:15
Felix Jahn - "Customizable Causal Enhancement of Reinforcement Learning Environments for Systematic Benchmarking".
Kunchangtai (Rickey) Liang - "DeDiCa: A Scalable Deconfounding Discrete Causal Generative Model".
Gerrit Großmann - "Rethinking Counterfactuals: Hidden Assumptions and Practical Pitfalls".
Lunch at Mensa 12:15 - 13:30
13:30 - 15:00
15:00 - 16:00
Dinner at Brauhaus zum Stiefel 6:00 - 9:00
Day 2
9:15 - 10:30
Abstract:
Causal reasoning sits at the heart of scientific discovery and remains one of the most demanding frontiers for AI. While large language models have demonstrated remarkable breadth across language tasks, their capacity for rigorous causal inference remains poorly understood, largely because the inferential pipeline from raw data to a credible causal claim demands tightly coupled decisions spanning structural assumption encoding, identifiability analysis, estimator selection, and sensitivity analysis. We argue that monolithic LLM prompting systematically fails to compose these steps with the required methodological discipline, and that decomposing causal inference into verifiable, agent-specialized subtasks is the right inductive bias for building AI systems that can reason causally in science.
This talk presents a unified research program that develops and stress-tests this hypothesis across the full identification-estimation pipeline. We begin by identifying where current agents fail, introducing a rigorous benchmark grounded in causal tasks from the published scientific literature that exposes systematic breakdowns in identification and estimator selection under realistic confounding regimes (CauSciBench). From this diagnostic foundation, we build toward autonomous causal inference with an end-to-end agent that takes observational data, metadata, and a causal query and performs covariate selection, backdoor and frontdoor identification, and effect estimation with uncertainty quantification (Causal AI Scientist).
We introduce other extensions to address the hardest subtasks in this pipeline: for settings requiring exogenous variation, we cast instrumental variable discovery as a multi-agent deliberation problem over domain knowledge and testable exclusion-restriction proxies (IV Co-Scientist); for verification, we introduce a symbolic layer that grounds LLM-produced causal claims in the do-calculus, catching identifiability violations and algebraic inconsistencies that chain-of-thought reasoning misses entirely (DoVerifier); for finding datasets that can answer a causal query, we introduce a retrieval system that matches on documented variables rather than titles or metadata (Causal Data Agent). Together, these systems define both the promise and the precise boundaries of what LLM-based causal agents can and cannot yet do in scientific discovery.
Finally, beyond this causal inference pipeline, we introduce CausalTutor, an interactive educational and research tool that enables practitioners to learn about causality and apply causal methods to their own domains.
Suggested Reading:
Causal AI Scientist: Facilitating Causal Data Science with Large Language Models. Paper Link
CauSciBench: Evaluating LLM Causal Inference for Scientific Research. https://icml.cc/virtual/2026/poster/60944
IV Co-Scientist: Multi-Agent LLM Framework for Causal Instrumental Variable Discovery (CLeaR 2026). arxiv.org/abs/2602.07943
Uncovering Hidden Correctness in LLM Causal Reasoning via Symbolic Verification (EACL 2026 Oral). arxiv.org/abs/2601.21210
Causal Tutor: An Interactive Research and Learning Tool for Causal Inference. Paper Link
Causal Data Agent: Variable-Aware Retrieval and Benchmarking for Causal Dataset Discovery. Paper Link. Paper Link
Quriosity: Analyzing Human Questioning Behavior and the Quest for Causality. (IJCNLP-AACL 2025 Findings). arxiv.org/abs/2405.20318
CLadder: Assessing Causal Reasoning in Language Models. (NeurIPS 2023). arxiv.org/abs/2312.04350
Can Large Language Models Infer Causation from Correlation? (ICLR 2024). arxiv.org/pdf/2306.05836
Beyond Memorization: Reasoning-Driven Synthesis as a Mitigation Strategy Against Benchmark Contamination. (ACL 2026)
Can Theoretical Physics Research Benefit from Language Agents? (NeurIPS 2025 AI4Science Workshop) https://arxiv.org/abs/2506.06214
Collective Intelligence: A Survey on Multi-Agent Systems for AI-Driven Scientific Discovery. https://www.preprints.org/manuscript/202508.1640/v1
Cooperate or Collapse: Emergence of Sustainable Cooperation in a Society of LLM Agents. (NeurIPS 2024). https://arxiv.org/abs/2404.16698
10:30 - 10:45
10:45 - 12:15
Christopher Lohse - "PRIM: Meta-Learned Bayesian Root Cause Analysis".
Hendrik Suhr - "Root Cause Analysis of Measurement and Mechanistic Anomalies".
Deborah Kanubala and Ayan Majumdar - “A Causal Framework to Measure and Mitigate Non-binary Treatment Discrimination”.
Lunch at Mensa 12:15 - 13:30
13:30 - 16:00
16:00 - 16:15