Our data and code are available at Agent Framework.
abstract
Large language model (LLM) agents are increasingly built on agent frameworks that provide reusable abstractions for workflow orchestration, state management, tool integration, and execution control. However, the quality of this infrastructure layer remains insufficiently understood, particularly its functionality challenges and usability concerns, as existing studies have mainly examined traditional deep learning (DL) frameworks or model-level agent failures. Therefore, we conduct an empirical study of 5,669 bug reports and 809 feature requests from five mainstream agent frameworks: AutoGen, CrewAI, LangChain, LangGraph, MetaGPT. We construct a four-dimensional taxonomy covering 22 root causes, seven symptoms, 11 motivations, and six requirements, and map them to the five-stage agent lifecycle.
Across the four RQs, results show an execution centered quality pattern shaped by semantic interface boundaries. Reported bugs mainly manifest as Incorrect Functionality (76.00%) and involve more API, configuration, parsing, and serialization related causes than DL framework bugs, while their associations remain sparse and stage specific. Feature requests mainly target Feature Enhancement (49.07%) and reveal structured needs for Orchestration Expressiveness, Development Delivery, Model Adaptation, and Tool Ecosystem. These findings call for quality assurance beyond crash based and tensor level testing, with emphasis on API sequences, structured LLM outputs, serialization boundaries, execution traces, and execution centered maintenance, offering empirical guidance for reliable and usable agent framework infrastructure.
Workflow of Our Four-phase Empirical Methodology
To investigate functional defects and usability concerns in modern Agent frameworks, we design a large-scale empirical study aiming at answering the following research questions:
• RQ1 (Agent Framework Failures): How do bug reports in agent frameworks distribute across root causes, symptoms, and lifecycle stages?
Finding : Agent framework bugs shift from low-level algorithmic and numerical faults toward semantic interface faults. This suggests that agent framework testing should prioritize API sequence testing, structured output mutation, schema validation, and serialization boundary checking.
Finding : Agent framework functionality challenges often take the form of silent semantic failures with weak crash signals. The decline of Crash from 51.40% in DL frameworks to 5.50% in agent frameworks suggests that crash-focused test oracles alone cannot effectively detect agent framework bugs. Agent framework testing therefore needs semantic oracles, such as state trajectory validation, expected output checking, tool call trace inspection, and execution trace replay.
Finding : Reported failures concentrate in the broad Self-Action stage. This concentration does not indicate that Self-Action has the highest normalized failure risk. Instead, it shows that many framework issues become visible in the central execution loop. This result motivates finer grained analysis of planning, tool invocation, memory and state handling, output parsing, execution control, and task result validation.
• RQ2 (Functionality Challenges Associations): What associations exist among root causes, symptoms, and lifecycle stages in agent frameworks?
Finding : Reported cause and symptom associations are sparse, while lifecycle stages show distinct bug signatures. These patterns call for association-aware and stage-aware testing, including configuration validation in Agent Initialization, input checks in Perception, concurrency testing in Mutual Interaction, and trace-based validation in Self-Action.
• RQ3 (Agent Framework Developer Evolution Requests): How do feature requests in agent frameworks distribute across motivations, requirements, and lifecycle stages?
Finding : Feature request motivations concentrate on Orchestration Expressiveness, Development Delivery, Model Adaptation, and Tool Ecosystem. This pattern shows a demand profile centered on orchestratable, deployable, adaptable, and integrable agent frameworks.
Finding : The dominance of Feature Enhancement (49.07%) over Feature Proposal (22.25%) shows a refinement oriented demand profile. Maintainers should prioritize robustness, flexibility, and developer convenience in existing features.
Finding : Feature requests concentrate on Self-Action (64.52%) but cover more lifecycle stages than functionality challenges. This pattern suggests that framework maintenance should strengthen the central execution loop while also addressing multi-agent coordination, initialization support, and long term adaptation.
• RQ4 (Usability Concerns Associations): What associations exist among motivations, requirements, and lifecycle stages in agent frameworks?
Finding : Developer facing requests exhibit structured associations among motivations, requirements, and lifecycle stages. Development Delivery aligns with documentation and infrastructure needs, while Model Adaptation aligns with integration needs. These patterns suggest that framework maintenance should organize feature requests by motivation, requirement type, and lifecycle context rather than treating them as isolated backlog items.
Due to the limited content of the text, further details are provided on the following page.
• root cause category
• symptom category
• motivation category
• requirement category
•result