Abstract: Large Language Models (LLMs) have revolutionized research across many domains, including knowledge engineering. However, ontology development remains a complex and labor-intensive task that requires substantial domain expertise to ensure adequate competency question coverage, adherence to modeling best practices, effective vocabulary reuse, and high-quality documentation. In this talk, Daniel will present an overview of their recent work on using LLMs and agentic AI to support ontology engineering, with a particular focus on the end-to-end generation of ontologies from a set of competency questions. They explore how LLM-based agents can leverage external tools to iteratively design, validate, refine, and document ontologies rather than simply generating them in a single pass. The talk will discuss the practical challenges of evaluating ontology quality, challenges of generating reliable benchmarks, as well as the importance of explainability, transparency, and traceability in the resulting ontologies. Overall, the talk will reflect on both the opportunities and limitations of LLMs in ontology engineering, and ask a broader question: should LLMs merely assist ontology engineers, or can they reliably take on a more autonomous role in the ontology engineering process?
Bio: Daniel Garijo is an Associate Professor in the Ontology Engineering Group at the Universidad Politécnica de Madrid (UPM). Previously, he worked as a researcher at the Information Sciences Institute of the University of Southern California in Los Angeles, where he also completed a postdoctoral fellowship between 2016 and 2018. Daniel's research focuses on Knowledge Engineering, Open Science and Semantic e-Science, particularly on capturing the context and metadata of scientific software and computational experiments to promote their reproducibility and reuse. To this end, he has proposed methods for creating vocabularies and knowledge graphs with data from computational experiments, exposed as objects accessible on the Web. Daniel has also contributed tools and best practices to assist knowledge engineers in the ontology engineering lifecycle (e.g., documentation, versioning and publication). At the international scene, Daniel has participated in W3C standardization groups (W3C Provenance Working Group), and is currently involved in the CodeMeta community to define software metadata standards, the consortium of Scientific repositories and registries (SciCodes) which brings together more than two dozen platforms for describing scientific software. In addition, he actively participates in European Open Science Cloud groups for FAIR and metrics and is a member of the Research Data Alliance group on FAIR and Machine Learning (FAIR4ML), which he co-leads. Daniel has participated in more than 20 national and international research projects, successfully carried out for agencies such as DARPA and NSF in the US and the European Commission in Europe. The standards he has helped develop have impacted hundreds of researchers in industry and academia.
09:00 - 09:10
🎊 Welcome & Workshop Opening Session
09:10 - 09:45
[🎙️Keynote] How good are Large Language Models at Ontology Engineering?
Dr. Daniel Garijo, Universidad Politécnica de Madrid (UPM)
09:45 - 10:00
[📃Presentation] OntoLLMJudge: A Framework for Neurosymbolic Evaluation of LLM Generations
Barbara Gendron, Stefani Tsaneva, Marta Sabou
10:00 - 10:15
[📃Presentation] The Impact of Textual Descriptions on Taxonomy Expansion with Large Language Models
Anastasiya Saputo, Artem Revenko
10:15 - 10:30
[📃Presentation] How Much Structure Is Enough? Evaluating Prompt Structuring Strategies for Knowledge-Constrained Configuration Migration
Massimiliano Fadda, Antonio D'Ambrosio, Diego Reforgiato Recupero, Paolo Platter
10:30 - 10:45
[📃Presentation] It Never Hurts to Ask: Evaluating LLMs for Ontology Metadata Enrichment
Davide Di Pierro, Danai Symeonidou, Lylia Abrouk
10:45 - 11:20
☕ Coffee Break & 📌 Poster Session
10:45 - 10:50
[📌Poster Pitch] When Does Memory Help an LLM Judge? A Leakage-Free, Audited Evaluation of Memory-Augmented Judging
Khush Patel, Bader Rasheed
10:50 - 10:55
[📌Poster Pitch] When Automatic Prompt Optimization Stops Helping: Evaluating TextGrad for Deterministic Policy–Descriptor Compliance Checking
Francesco Simbola, Martina Salis, Diego Reforgiato Recupero, Daniele Riboni
10:55 - 11:00
[📌Poster Pitch] Toward Safe Healthcare LLMs: Alignment Faking, Introspection, and Control
Antonello Meloni, Antonio Puertas Gallardo, Sergio Consoli, Lorenzo Bertolini, Mario Ceresa, Diego Reforgiato Recupero
11:00 - 11:20
☕ Poster Viewing, Coffee & Discussion
Participants can visit the posters and discuss the work directly with the authors.
11:20 - 11:35
[📃Presentation] Events Matter: Delegating instead of Training
Erwin Filtz, María Navas-Loro, Sabrina Kirrane
11:35 - 11:50
[📃Presentation] rudofMCP: A Model Context Protocol Server Exposing Knowledge Graph Engineering Capabilities to Large Language Models
Samuel Bustamante Larriet, Jose Emilio Labra-Gayo, Álvaro García-Fernández
11:50 - 12:05
[📃Presentation] The LocalEvents Benchmark: Event Extraction and Slot Filling with Small LLMs
Smirnou Uladizslau, Adrian M.P. Brasoveanu, Lyndon J.B. Nixon
12:05 - 12:20
[📃Presentation] Evaluating Ontology-Constrained LLM Correction in Scientific Text
Manasi Gokhale, Marta Dembska
12:20 - 12:35
💬 Discussion & Feedback Session
12:35 - 12:50
🎊 Closing Session & Networking
Sala Cigno,
The Nicolaus Hotel,
Bari,
Italy