Agentic AI systems increasingly operate across multiple tools, APIs, and data sources but existing benchmarks evaluate these capabilities in isolation. This SemEval 2027 task introduces VAKRA Advanced, a challenging benchmark designed to evaluate multi-hop, multi-source reasoning in realistic, tool-grounded environments. Participants will develop systems that can plan, select, and chain API calls, while integrating information from both structured tools and unstructured documents. Through four focused sub-tasks, VAKRA Advanced provides a comprehensive framework for assessing the next generation of intelligent, tool-using agents.
Sub Task 1: API Chaining
Sub Task 2: Tool Selection
Sub Task 3: Multi-hop API Reasoning
Sub Task 4: Multi-hop, Multi-source Reasoning
Ankita Rajaram Naik , IBM
Anupama Murthi, IBM
Benjamin Elder,
Siyu Huo, IBM
Praveen Venkateswaran, IBM
Hamid Adebayo, IBM
Sara Rosenthal, IBM
Danish Contractor, IBM
Contact: ankita.naik@ibm.com
Sample data and environment release - Checkout the environment setup [GitHub] and sample data [HuggingFace Dataset]
Training Data Release - Training dataset is uploaded to [HuggingFace Dataset]
Evaluation Start Date - TBD
Evaluation End Date - TBD
Paper submission due February 2027
Notification to authors March 2027
Camera ready due April 2027