In this section, we primarily analyze the root causes of functional bugs in Agent Frameworks.
In our study, focusing on five mainstream Agent Frameworks—namely AutoGen, CrewAI, LangChain, LangGraph, and MetaGPT—we screened 5,669 bug reports from GitHub issues and categorized the root causes of these bugs. As a result, we constructed 22 categories of root causes to explain the mechanisms underlying different types of defects in Agent Frameworks.
In addition, we also compare the root cause of the Agent Framework with that of the DL framework (Compared with DL Framework).
C1. API Incompatibility
The kind of root cause in agent frameworks refers to changes in function signatures, object patterns, or invocation protocols of framework-internal abstractions that break existing agent configurations and pipelines.
For example:
In LangGraph5830, due to changes in the structure of the Interrupt object in LangGraph 0.6+, the legacy version langgraph-api==0.2.100 still accesses the `interrupt.resumable` property according to the old interface.
C2. API Misuse
The kind of root cause in agent frameworks arises when developers misinterpret the framework’s execution abstractions, control flow semantics, or environmental assumptions.
For example:
In CrewAI1697, the developer misunderstood the usage of the CrewAI `@crew` decorator and directly invoked the uninitialized internal attribute `self.crewer`, instead of creating or obtaining a Crew instance through the framework-provided `self.crew()` method.
C3. Misconfiguration
The kind of root cause in agent frameworks stems from discrepancies between the runtime configuration expected by the framework and the actual environment.
For example:
When using OpenRouter as an OpenAI-compatible interface in AutoGen6502, the model name and model_info configuration do not conform to the requirements of AutoGen/OpenAIChatCompletionClient.
C4. Telemetry Malfunction
The kind of root cause in agent frameworks involves monitoring, tracing, or observability mechanisms that interfere with or fail during execution, interrupting the main workflow due to strict validation, intrusive exception handling, or network certificate issues.
For example:
After enabling `langchain.debug = True` in LangGraph4119, the framework's `ConsoleCallbackHandler` fails to locate the corresponding run ID during streaming token callbacks.
C5. Concurrency Issue
The kind of root cause in agent frameworks manifests as inconsistent or prematurely terminated asynchronous event streams coordinating model outputs, tool responses, or execution signals, or as misunderstandings of the concurrency model.
For example:
In AutoGen5563, within the Swarm, the speaking order, handoff timing, and duplicate utterance control among multiple agents are not stably constrained, resulting in a situation where an agent may either refrain from sending any message before a handoff or send multiple duplicate messages consecutively.
C6. Serialization Error
The kind of root cause in agent frameworks refers to faults in encoding or decoding structured execution data.
For example:
In LangChain8858, the LLMChain can be invoked normally before saving; however, after save() and load_chain(), the internal HuggingFacePipeline.pipeline is not properly deserialized and restored, becoming None.
C7. Memory Persistence Bug
The kind of root cause in agent frameworks occurs when mechanisms for maintaining long-term agent state, embedding caches, or retrieval indices fail to remain consistent with evolving knowledge sources or model configurations due to design flaws.
For example:
In LangGraph5604, when saving and reading checkpoint metadata, the checkpoint-sqlite implementation handles thread persistence fields such as `thread_id` and `thread_ts` inconsistently, resulting in persisted metadata that does not match the test expectations.
C8. Knowledge Transmission Bug
The kind of root cause in agent frameworks appears in hierarchical or multi-agent architectures as defects in context transfer logic.
For examle:
In CrewAI1703, although knowledge has been preloaded into the crew/agent, the planning agent does not receive this knowledge when generating plans.
C9. Console Interaction Bug
The kind of root cause arises when the agent framework’s assumptions about interactive execution environments conflict with the actual runtime context.
For example:
In LangChain10494, the human tool originally relied on interactive human input, but after migrating from Jupyter to Streamlit, the input function `get_input` is not triggered.
C10. Frontend–Backend Mismatch
The kind of root cause in agent frameworks occurs in agent frameworks that provide a web interface, where defects in frontend–backend state synchronization logic cause the frontend to display outdated or inconsistent information.
For example:
In AutoGen5872, the edit button in the AutoGen Studio frontend UI is not properly bound to the editing logic of the termination condition, resulting in no update of the frontend state after the user clicks it.
C11. Documentation Desync
The kind of root cause in agent frameworks reflects inconsistencies between the framework documentation and its actual implementation.
For example:
In LangChain6332, the tutorial still uses the legacy import methods `langchainhub` / `FunctionMessage`, whereas the latest version of LangChain has migrated to `langchain_classic.hub` / `ToolMessage`.
C12. Model Output Parsing Error
The kind of root cause in agent frameworks arises when content returned by the language model deviates from the expected format, length, or contains invalid characters, preventing the parser from extracting valid information.
For example:
In MetaGPT1050, the user input is "meat vs icecream", but gemini-pro returns irrelevant keywords such as "Computer Vision / Object Detection".
C13. Resource Limitation
The kind of root cause in agent frameworks involves constraints imposed by resource caps, such as token limits, memory, time, or concurrency, or improper resource management.
For example:
In LangChain7827, during vectorization/insertion, the number of texts passed in a single batch exceeds the embedding service's batch size limit of ≤ 10.
C14. Dependent Module Issue
The kind of root cause in agent frameworks occurs when external services return errors, time out, or impose rate limits, or when the model itself lacks support for a required feature and the framework fails to handle such scenarios.
For example:
The core issue in MetaGPT1059 is that when using a local LLM service (LMStudio + HuggingFace model), the model responds too slowly, resulting in an `httpcore.ReadTimeout`.
C15. Incorrect Algorithm Implementation
The kind of root cause in agent frameworks refers to bugs at the implementation level of the framework’s own code.
For example:
In LangGraph5182, when rendering nodes with `defer=True`, the framework's algorithm for inferring node execution order and the END edge is flawed.
C16. Environment Incompatibility
The kind of root cause refers to errors caused by overlooking specific characteristics of the execution environment, such as hardware platforms or operating systems.
For example:
In CrewAI2409, the `chroma-hnswlib` dependency required by `crewai[tools]` needs local compilation on Windows ARM architecture, but the build process fails to correctly link the MSVC libraries, resulting in a wheel build failure. This issue occurs only on ARM Windows devices, whereas older Intel Windows computers are not affected.
C17. Incorrect Exception Handling
The kind of root cause refers to defects in exception handling, including missing exceptions that should be thrown or handled, redundant exceptions that should not be thrown, and incorrect or imprecise exception messages.
For example:
In LangChain8711, the framework catches the actual missing dependency exception but throws an ambiguous error message, concealing the root cause that `nltk` is not installed.
C18. Type Issue
The kind of root cause involves type-related problems, such as errors in type conversion or type checking.
For example:
In LangGraph6207, the framework's type definition is too narrow, causing a valid list of BaseMessage to be erroneously rejected during static type checking.
C19. Tensor Shape Mismatch
The kind of root cause occurs when tensor shapes are incompatible in shape-related operations, such as shape inference or tensor transformation.
For example:
In LangChain15730, the PDF image object declares dimensions of 193 × 121, but the actual array size read is only 293, which cannot be reshaped to the corresponding shape.
C20. Incorrect Assignment
The kind of root cause refers to errors caused by incorrectly assigning values to variables or by missing variable initialization.
For example:
In AutoGen5186, a typographical omission of the letter 'n' in "ConectionStrings" causes the variable/configuration item to be assigned to an incorrect key, resulting in the framework failing to locate the value when reading "ConnectionStrings:HelloAIAgents".
C21. Numerical Issue
The kind of root cause is caused by incorrect numerical computation, such as division by zero, overflow or underflow, incorrect operators or operands, or missing operands.
For example:
In LangChain24698, the retriever returns negative similarity scores under the `similarity_score_threshold` mode.
C22. Others
The kind of root cause in agent frameworks encompasses issues that do not fit into any of the previous twenty-one categories.