option
Home
News
Why Agents Forget and Go Off-Track in Long Tasks: AWS, Claude Code, and Manus Unpack Four Frameworks

Why Agents Forget and Go Off-Track in Long Tasks: AWS, Claude Code, and Manus Unpack Four Frameworks

October 2, 2026
3

Why Agents Forget and Go Off-Track in Long Tasks: AWS, Claude Code, and Manus Unpack Four Frameworks

Large language models frequently lose focus during extended operations, drifting away from their original objectives—a persistent challenge in the intelligent agent sector. The issue often stems not from the model itself, but from the surrounding execution framework. A recent AWS design guide for cloud programming agents highlights this clearly: shallow agents suffer from context overflow, lose focus during long cycles, and fail to maintain state. Addressing this requires optimizing the harness, which manages all operations external to the core model.

A comprehensive review of leading frameworks has shed light on this critical layer. LangChain Deep Agents, Claude Code, Manus, OpenAI Codex, and Amazon Bedrock AgentCore each utilize four key mechanisms—context budgeting, compression, task state management, and cross-session memory—to convert basic loops into robust agents capable of handling complex, long-running tasks.

The most counterintuitive finding is that simply expanding the context window does not resolve the issue. Chroma’s Context Rot report, which evaluated 18 large models, revealed that even in straightforward retrieval tasks, model reliability decreases as input length increases. Anthropic explains that attention mechanisms generate quadratic pairwise relationships for n tokens, meaning each additional token consumes a finite portion of the attention budget. Context is a depleting resource, not an infinite container. Manus further noted that typical tasks involve approximately 50 tool calls, with an input-to-output ratio nearing 100:1. Consequently, initial instructions gradually shift toward the center of the window—precisely where memory degradation is most pronounced.

The first engine focuses on context budgeting and offloading. Deep Agents enforce two strict rules: if a tool returns more than 20,000 tokens, the data is written to the file system, retaining only the file path and a 10-line preview; when session context exceeds 85% of the window capacity, older edit commands are replaced with pointers. Claude Code applies similar logic before loading, capping memory usage at 200 lines or 25KB, with MCP tool mode defaulting to listing names and fetching details on demand. AWS AgentCore’s implementation is even more rigorous: the coordinator spawns three browser sub-agents in parallel, each operating within its own MicroVM. The analysis sub-agent receives only structured results, reducing expected duration to 4–6 minutes, whereas serial execution can take up to three times longer.

The second engine handles compression. When offloading is insufficient, the framework summarizes the conversation near capacity limits and restarts the process. Claude Code’s compression prompt preserves architectural decisions and unresolved bugs while discarding redundant outputs. After compression, it re-reads the last five modified files and re-injects relevant skill text. It also archives the complete original record to disk, ensuring that facts lost during summarization can be retrieved later. Deep Agents treats goal preservation as a structural feature, organizing summary documents into intent, generated outputs, and next steps. Compression has also been integrated at the API level: OpenAI’s Responses API offers server-side compression via `compact_threshold`, which Codex leverages for long programming tasks. The Claude platform provides customizable compression settings with writable instructions.

The third engine manages task state and persistent reminders. Manus employs a straightforward approach: creating a `todo.md` file and checking off items step-by-step, effectively embedding the goal at the end of the context to counteract being “submerged in the middle.” However, this method is not universally effective: Deep Agents made its to-do middleware optional in version 0.7 (released July 2026), as evaluations showed that disabling it slightly improved reward scores and reduced costs. LangChain still recommends re-enabling this feature for long tasks, weaker models, and interfaces requiring progress visibility.

The fourth engine enables cross-session memory. Claude Code reloads CLAUDE.md and automatic memory after each compression cycle. AgentCore Memory runs extraction strategies in the background, allowing the coordinator to recall previous insights directly rather than re-analyzing them. However, research from ETH Zurich suggests caution: context files like AGENTS.md typically do not improve success rates but increase reasoning costs by 20% to 23%, imposing a fixed tax on the attention budget with each reload. Therefore, Claude Code advises keeping CLAUDE.md under 200 lines.

The study concludes that verifying whether a framework truly “grasps the goal” requires forced compression tests. The most dangerous failure mode occurs when an agent immediately requests clarification after a summary or incorrectly declares a task complete. Ultimately, keeping an agent on track during long journeys depends not on the size of the model, but on the precision of these supporting engines.

Related article
Qwen AI Platform Expands Model Matrix With Official Launch of GLM-5.3 and DeepSeek-V4-Pro Qwen AI Platform Expands Model Matrix With Official Launch of GLM-5.3 and DeepSeek-V4-Pro The Alibaba Cloud Qwen AI platform (MaaS) has recently expanded its model service matrix, integrating Zhipu’s flagship large model GLM-5.3 and the official release of DeepSeek-V4-Pro. The corresponding API services are now publicly available, enablin
China’s AI sector sees full-chain breakthroughs, accelerating Artificial Intelligence Law China’s AI sector sees full-chain breakthroughs, accelerating Artificial Intelligence Law Global downloads of domestic large language models have surpassed 10 billion, with trillion-parameter open-source models emerging regularly. China’s artificial intelligence sector is achieving comprehensive breakthroughs across its entire value chain
Zhiyuan Innovation Unveils Data Acquisition 2.0 to Power Smarter Robots Zhiyuan Innovation Unveils Data Acquisition 2.0 to Power Smarter Robots As the artificial intelligence sector expands rapidly, enabling robots to accurately comprehend and adapt to intricate real-world settings has emerged as a critical industry priority. During the Tianfu Artificial Intelligence Industry Ecosystem and P
Related Special Topic Recommendations
Health & Wellness ChatGPT Meal Planning Apps for Macro Tracking, Grocery Lists, and Weekly Prep
ChatGPT Meal Planning Apps for Macro Tracking, Grocery Lists, and Weekly Prep

2026 Latest Best ChatGPT Meal Planning Apps Top-rated must-try curated powerful game-changing tools for macro tracking grocery lists and weekly prep. XIX.AI delivers weekly updated rankings based on real-world tests, including free vs paid comparison details. Discover your perfect tool to boost productivity and unlock your AI edge today. Explore now!

9 tools
xix.ai
Video creation AI Video Script Tools for Short Form Content
AI Video Script Tools for Short Form Content

2026 Latest Best Top-Rated AI Video Script Tools for Short Form Content! XIX.AI has curated a powerful, game-changing collection of must-try tools that go through rigorous real-world tests. You’ll find detailed weekly updated rankings, free vs paid comparison info, and insights to help you boost productivity and unlock your AI edge in creating high-quality short form content. Explore now to discover your perfect tool!

10 tools
xix.ai
Comic Creation AI Comic Panel Tools for Storyboard Style Narratives
AI Comic Panel Tools for Storyboard Style Narratives

2026 Latest Best AI Comic Panel Tools for Storyboard Style Narratives! This curated top-rated collection features powerful game-changing solutions that help creators boost writing efficiency and overcome creation bottlenecks. You’ll find detailed free vs paid comparison, real-world tests, and official rankings to guide your pick. XIX.AI offers expert insights too. Explore now to Discover your perfect tool for bringing vivid stories to life! 238 characters

9 tools
xix.ai
Video creation Best AI Video Generation Tools for Short Content
Best AI Video Generation Tools for Short Content

2026 Latest Best Top-rated AI Video Generation Tools for Short Content are here on XIX.AI’s curated collection. These powerful game-changing tools have undergone rigorous real-world tests, featuring a comprehensive free vs paid comparison and detailed rankings to help you find the perfect option. Whether you need to create high-quality short videos fast or boost your productivity across various tasks, these must-try solutions deliver outstanding results. Explore now to discover your ideal tool and unlock your AI edge.

9 tools
xix.ai
Design & Art Best AI Art Generators: Explore Moodboards, Styles, and Visual Directions
Best AI Art Generators: Explore Moodboards, Styles, and Visual Directions

2026 Latest Best AI Art Generators ranking showcases top-rated tools for creating stunning moodboards, exploring diverse styles, and defining visual directions. XIX.AI offers a curated list of powerful, game-changing options that have passed rigorous real-world tests. Get a free vs paid comparison to help you find the perfect fit for boosting your creative workflow. Explore now to Unlock your AI edge in art creation.

10 tools
xix.ai
Marketing AI Email Personalization Platforms for Shopify Flows, Retention, and Win-Backs
AI Email Personalization Platforms for Shopify Flows, Retention, and Win-Backs

2026 Latest Best AI Email Personalization Platforms for Shopify Flows that boost retention and win-backs. This top-rated curated list features powerful game-changing tools经过real-world tests, offering free vs paid comparison insights. XIX.AI is part of our trusted selection. Discover your perfect tool to unlock your AI edge today. Explore now!

9 tools
xix.ai
Comments (0)
0/500
OR