Home
Why Agents Forget and Go Off-Track in Long Tasks: AWS, Claude Code, and Manus Unpack Four Frameworks

Large language models frequently lose focus during extended operations, drifting away from their original objectives—a persistent challenge in the intelligent agent sector. The issue often stems not from the model itself, but from the surrounding execution framework. A recent AWS design guide for cloud programming agents highlights this clearly: shallow agents suffer from context overflow, lose focus during long cycles, and fail to maintain state. Addressing this requires optimizing the harness, which manages all operations external to the core model.
A comprehensive review of leading frameworks has shed light on this critical layer. LangChain Deep Agents, Claude Code, Manus, OpenAI Codex, and Amazon Bedrock AgentCore each utilize four key mechanisms—context budgeting, compression, task state management, and cross-session memory—to convert basic loops into robust agents capable of handling complex, long-running tasks.
The most counterintuitive finding is that simply expanding the context window does not resolve the issue. Chroma’s Context Rot report, which evaluated 18 large models, revealed that even in straightforward retrieval tasks, model reliability decreases as input length increases. Anthropic explains that attention mechanisms generate quadratic pairwise relationships for n tokens, meaning each additional token consumes a finite portion of the attention budget. Context is a depleting resource, not an infinite container. Manus further noted that typical tasks involve approximately 50 tool calls, with an input-to-output ratio nearing 100:1. Consequently, initial instructions gradually shift toward the center of the window—precisely where memory degradation is most pronounced.
The first engine focuses on context budgeting and offloading. Deep Agents enforce two strict rules: if a tool returns more than 20,000 tokens, the data is written to the file system, retaining only the file path and a 10-line preview; when session context exceeds 85% of the window capacity, older edit commands are replaced with pointers. Claude Code applies similar logic before loading, capping memory usage at 200 lines or 25KB, with MCP tool mode defaulting to listing names and fetching details on demand. AWS AgentCore’s implementation is even more rigorous: the coordinator spawns three browser sub-agents in parallel, each operating within its own MicroVM. The analysis sub-agent receives only structured results, reducing expected duration to 4–6 minutes, whereas serial execution can take up to three times longer.
The second engine handles compression. When offloading is insufficient, the framework summarizes the conversation near capacity limits and restarts the process. Claude Code’s compression prompt preserves architectural decisions and unresolved bugs while discarding redundant outputs. After compression, it re-reads the last five modified files and re-injects relevant skill text. It also archives the complete original record to disk, ensuring that facts lost during summarization can be retrieved later. Deep Agents treats goal preservation as a structural feature, organizing summary documents into intent, generated outputs, and next steps. Compression has also been integrated at the API level: OpenAI’s Responses API offers server-side compression via `compact_threshold`, which Codex leverages for long programming tasks. The Claude platform provides customizable compression settings with writable instructions.
The third engine manages task state and persistent reminders. Manus employs a straightforward approach: creating a `todo.md` file and checking off items step-by-step, effectively embedding the goal at the end of the context to counteract being “submerged in the middle.” However, this method is not universally effective: Deep Agents made its to-do middleware optional in version 0.7 (released July 2026), as evaluations showed that disabling it slightly improved reward scores and reduced costs. LangChain still recommends re-enabling this feature for long tasks, weaker models, and interfaces requiring progress visibility.
The fourth engine enables cross-session memory. Claude Code reloads CLAUDE.md and automatic memory after each compression cycle. AgentCore Memory runs extraction strategies in the background, allowing the coordinator to recall previous insights directly rather than re-analyzing them. However, research from ETH Zurich suggests caution: context files like AGENTS.md typically do not improve success rates but increase reasoning costs by 20% to 23%, imposing a fixed tax on the attention budget with each reload. Therefore, Claude Code advises keeping CLAUDE.md under 200 lines.
The study concludes that verifying whether a framework truly “grasps the goal” requires forced compression tests. The most dangerous failure mode occurs when an agent immediately requests clarification after a summary or incorrectly declares a task complete. Ultimately, keeping an agent on track during long journeys depends not on the size of the model, but on the precision of these supporting engines.
Related article
Qwen AI Platform Expands Model Matrix With Official Launch of GLM-5.3 and DeepSeek-V4-Pro
The Alibaba Cloud Qwen AI platform (MaaS) has recently expanded its model service matrix, integrating Zhipu’s flagship large model GLM-5.3 and the official release of DeepSeek-V4-Pro. The corresponding API services are now publicly available, enablin
China’s AI sector sees full-chain breakthroughs, accelerating Artificial Intelligence Law
Global downloads of domestic large language models have surpassed 10 billion, with trillion-parameter open-source models emerging regularly. China’s artificial intelligence sector is achieving comprehensive breakthroughs across its entire value chain
Zhiyuan Innovation Unveils Data Acquisition 2.0 to Power Smarter Robots
As the artificial intelligence sector expands rapidly, enabling robots to accurately comprehend and adapt to intricate real-world settings has emerged as a critical industry priority. During the Tianfu Artificial Intelligence Industry Ecosystem and P
Related Special Topic Recommendations
Comments (0)
0/500

Large language models frequently lose focus during extended operations, drifting away from their original objectives—a persistent challenge in the intelligent agent sector. The issue often stems not from the model itself, but from the surrounding execution framework. A recent AWS design guide for cloud programming agents highlights this clearly: shallow agents suffer from context overflow, lose focus during long cycles, and fail to maintain state. Addressing this requires optimizing the harness, which manages all operations external to the core model.
A comprehensive review of leading frameworks has shed light on this critical layer. LangChain Deep Agents, Claude Code, Manus, OpenAI Codex, and Amazon Bedrock AgentCore each utilize four key mechanisms—context budgeting, compression, task state management, and cross-session memory—to convert basic loops into robust agents capable of handling complex, long-running tasks.
The most counterintuitive finding is that simply expanding the context window does not resolve the issue. Chroma’s Context Rot report, which evaluated 18 large models, revealed that even in straightforward retrieval tasks, model reliability decreases as input length increases. Anthropic explains that attention mechanisms generate quadratic pairwise relationships for n tokens, meaning each additional token consumes a finite portion of the attention budget. Context is a depleting resource, not an infinite container. Manus further noted that typical tasks involve approximately 50 tool calls, with an input-to-output ratio nearing 100:1. Consequently, initial instructions gradually shift toward the center of the window—precisely where memory degradation is most pronounced.
The first engine focuses on context budgeting and offloading. Deep Agents enforce two strict rules: if a tool returns more than 20,000 tokens, the data is written to the file system, retaining only the file path and a 10-line preview; when session context exceeds 85% of the window capacity, older edit commands are replaced with pointers. Claude Code applies similar logic before loading, capping memory usage at 200 lines or 25KB, with MCP tool mode defaulting to listing names and fetching details on demand. AWS AgentCore’s implementation is even more rigorous: the coordinator spawns three browser sub-agents in parallel, each operating within its own MicroVM. The analysis sub-agent receives only structured results, reducing expected duration to 4–6 minutes, whereas serial execution can take up to three times longer.
The second engine handles compression. When offloading is insufficient, the framework summarizes the conversation near capacity limits and restarts the process. Claude Code’s compression prompt preserves architectural decisions and unresolved bugs while discarding redundant outputs. After compression, it re-reads the last five modified files and re-injects relevant skill text. It also archives the complete original record to disk, ensuring that facts lost during summarization can be retrieved later. Deep Agents treats goal preservation as a structural feature, organizing summary documents into intent, generated outputs, and next steps. Compression has also been integrated at the API level: OpenAI’s Responses API offers server-side compression via `compact_threshold`, which Codex leverages for long programming tasks. The Claude platform provides customizable compression settings with writable instructions.
The third engine manages task state and persistent reminders. Manus employs a straightforward approach: creating a `todo.md` file and checking off items step-by-step, effectively embedding the goal at the end of the context to counteract being “submerged in the middle.” However, this method is not universally effective: Deep Agents made its to-do middleware optional in version 0.7 (released July 2026), as evaluations showed that disabling it slightly improved reward scores and reduced costs. LangChain still recommends re-enabling this feature for long tasks, weaker models, and interfaces requiring progress visibility.
The fourth engine enables cross-session memory. Claude Code reloads CLAUDE.md and automatic memory after each compression cycle. AgentCore Memory runs extraction strategies in the background, allowing the coordinator to recall previous insights directly rather than re-analyzing them. However, research from ETH Zurich suggests caution: context files like AGENTS.md typically do not improve success rates but increase reasoning costs by 20% to 23%, imposing a fixed tax on the attention budget with each reload. Therefore, Claude Code advises keeping CLAUDE.md under 200 lines.
The study concludes that verifying whether a framework truly “grasps the goal” requires forced compression tests. The most dangerous failure mode occurs when an agent immediately requests clarification after a summary or incorrectly declares a task complete. Ultimately, keeping an agent on track during long journeys depends not on the size of the model, but on the precision of these supporting engines.
Qwen AI Platform Expands Model Matrix With Official Launch of GLM-5.3 and DeepSeek-V4-Pro
The Alibaba Cloud Qwen AI platform (MaaS) has recently expanded its model service matrix, integrating Zhipu’s flagship large model GLM-5.3 and the official release of DeepSeek-V4-Pro. The corresponding API services are now publicly available, enablin
China’s AI sector sees full-chain breakthroughs, accelerating Artificial Intelligence Law
Global downloads of domestic large language models have surpassed 10 billion, with trillion-parameter open-source models emerging regularly. China’s artificial intelligence sector is achieving comprehensive breakthroughs across its entire value chain
Zhiyuan Innovation Unveils Data Acquisition 2.0 to Power Smarter Robots
As the artificial intelligence sector expands rapidly, enabling robots to accurately comprehend and adapt to intricate real-world settings has emerged as a critical industry priority. During the Tianfu Artificial Intelligence Industry Ecosystem and P











