VAKRA-Advanced evaluates agentic systems in executable environments where agents must reason across locally hosted APIs, databases, and document collections.
The released benchmark is organized into train and test splits for each subtask. Inputs are dialogue-style JSON instances. Train examples include expected answers and ground-truth tool-call sequences, while the official evaluation split uses held-out and newly curated instances.
This task tests whether an agent can plan and execute a sequence of API calls. It uses business-intelligence-style APIs with two interface designs: compositional interfaces, where agents combine generic operations such as filtering, aggregation, and transformation, and expanded function interfaces, where parameterized operations appear as distinct tools. The challenge is to compose the right steps and maintain intermediate state across calls.
This task focuses on choosing the right endpoint for the user request. Endpoint-style interfaces expose specific, query-aligned tools that encapsulate much of the computation. Because each call is more specialized, the main difficulty is accurate query interpretation and selecting the correct endpoint from the available tool set.
This task requires 3-5 connected reasoning steps over endpoint-style APIs. Outputs from earlier calls determine the inputs to later calls, so agents must extract parameters, disambiguate entities, and chain identifiers across semantically distinct endpoints while preserving the user's intent.
This task extends multi-hop API reasoning to settings that combine structured APIs with unstructured document collections. Agents must decide when to retrieve documents, what evidence to extract, how to align text-derived information with API parameters, and how to reconcile naming or format inconsistencies across sources in single-turn and multi-turn settings.