This tutorial introduces RecoAtlas, a benchmark and tool-based environment for evaluating LLM-based recommendation agents on product search and set-construction tasks. The focus is on moving beyond semantic plausibility toward set-level utility: can an agent retrieve, compare, and compose products into coherent and useful recommendation sets?
We will walk through the benchmark, the available tools, and a practical fine-tuning setup for building a domain-specialized shopping assistant, with outfit generation as the main running example. The tutorial will cover task formulation, tool use, retrieval, bundle construction, evaluation, and post-training directions for agents that reason over products and optimize for both semantic coherence and recommendation utility.
Flavian Vasile
Otmane Sakhi
Imad Aouali