// Journal · Sep 15, 2026 · 5 min read

Context engineering for AI agents: why yours forgets

What is context engineering for AI agents? It is the practice of deciding, at every step of a task, exactly which tokens the model sees — instructions, tool definitions, conversation history, retrieved documents — and ruthlessly cutting everything else. Prompt engineering asks “what should I tell the model?” Context engineering asks “what should be in the window right now?” In our experience, when a client says “the agent got dumber halfway through,” the model didn’t change. The context did.

That distinction became the working vocabulary of agent development after Anthropic’s engineering team published Effective context engineering for AI agents, and it matches what we see in the automations we build: the difference between an agent that finishes a forty-step task and one that wanders off at step twelve is almost never the model. It’s what was in the window at step twelve.

What context engineering actually means

An agent’s context window is everything the model can “see” when it produces its next action: the system prompt, the tools it’s allowed to call, the results of the tools it already called, and the history of the run so far. On a long task, that history grows every step — and the window is finite.

Context engineering is the discipline of treating those tokens as a budget. Every tool definition, every pasted document, every verbose tool result spends some of it. Anthropic’s framing is that models have a limited attention budget that depletes as the window fills — so the goal is the smallest set of high-signal tokens that makes the next step obvious, not the largest set of possibly-relevant ones.

This is why “just give it everything” fails as an agent strategy. More context is not more knowledge. Past a point, it’s noise the model has to attend past.

Context rot: agents degrade before the window is full

The forgetting has been measured. Chroma’s context rot research evaluated 18 frontier models — including GPT-4.1, Claude 4, and Gemini 2.5 — and found that every one of them gets less reliable as input length grows, even on tasks as simple as retrieving a stated fact or replicating text.

Two findings matter for anyone running agents in production:

  • Degradation starts well before the window limit. A model with a 200K-token window can be measurably worse at 50K tokens. Hitting the maximum is not the failure mode; the slow slide on the way there is.
  • Position matters. Models recall information at the beginning and end of the context far better than information buried in the middle — which is exactly where a fact from step 5 of a 40-step run ends up.

So an agent that “forgets” the customer’s account ID it looked up twenty steps ago isn’t malfunctioning. It’s behaving exactly as the research predicts. The fix isn’t a bigger window — it’s not letting the run accumulate that much low-value history in the first place.

Four techniques that keep long-running agents on track

These are the patterns Anthropic documents and the ones we reach for in practice:

  1. Compaction. When history gets long, summarise it and start a fresh window that keeps the decisions and discards the transcript. The critical detail is what survives the summary: open questions, IDs, constraints — not a prose recap.
  2. Structured note-taking. The agent writes durable facts to a scratchpad outside the window — a file, a database row — and reads them back when needed. Memory that doesn’t decay because it isn’t competing for attention.
  3. Sub-agents. A focused agent does the deep exploration — reading a 30-page document, crawling a codebase — and returns a condensed summary of 1,000–2,000 tokens to the main agent. The mess stays in the sub-agent’s window and dies with it.
  4. Just-in-time retrieval. Instead of front-loading every document that might matter, give the agent tools to fetch what it needs when it needs it. A file path costs a handful of tokens; the file costs thousands — load it only at the step that uses it.

None of these is exotic. All four are plumbing. That’s the general shape of production AI work: the model is the smallest part of the system.

Is it a context problem or a model problem?

A quick diagnostic we use before anyone blames the model:

  • Does it fail late but not early? Same task, works in the first ten steps, drifts after thirty — that’s context rot, not capability. A capability gap fails at step one too.
  • Does it fail on long inputs but pass on excerpts? Paste only the relevant section and it gets the answer right — the model was fine; the retrieval was lazy.
  • Does it ignore an instruction it followed earlier? The instruction is probably now mid-window, in the position models recall worst. Restate critical constraints near the end of context, or re-inject them each step.
  • Does it call the wrong tool from a set of thirty? Overlapping tool definitions are ambiguity the model has to resolve every single step. Cut the set down before touching the prompt.

If none of those match — the agent fails immediately, on short context, with clear instructions — then you may genuinely have a model or task-design problem. In our experience that’s the rare case.

What this means for business automation

Havoric is an AI automation and web development agency: we automate repetitive manual processes and build the web and mobile apps around them, and most of what makes those automations reliable is context discipline, not model choice. When we scope AI automation work, “which model?” is a ten-minute decision. What the agent sees at each step — which tools, which records, how history is compacted — is where the engineering weeks go.

It also pairs with the pattern we’ve written about before: a review queue that keeps humans in the loop. Context engineering raises the share of cases the agent handles correctly; the queue catches the rest. Neither replaces the other.

If your agent worked in the demo and forgets in production, don’t reach for a bigger model. Read the window. The answer is almost always in there — usually in the middle, where the model can’t see it.