The Problem: Context Rot Makes Large Context Windows a Liability
Context engineering—the discipline of deciding exactly which tokens enter an LLM's context window and when—has replaced prompt engineering as the primary bottleneck in agentic AI for software engineering. The assumption that a 1M-token context window solves agent memory problems is empirically wrong. 2025 research on context rot demonstrates that model recall and reasoning accuracy degrade non-linearly as token counts climb, even well inside nominal limits: relevant facts buried at position 400K are routinely missed at rates several times higher than the same facts at position 4K. For AI coding agents, this manifests as duplicated function creation, forgotten conventions mid-task, and regressions in files the agent already edited.
- Attention dilution: More tokens means more distractors competing with the actual task signal.
- Instruction drift: Agents progressively ignore system-prompt constraints as working context grows.
- Cache invalidation: Prompt caching (offering up to ~90% input-token discounts) is destroyed by mutating early tokens, so careless context assembly is also a cost and latency problem.
- Empirical confirmation: Benchmarks like SWE-bench repeatedly show agents resolving issues with tighter, more targeted context outperforming agents stuffed with entire repository dumps.
Anthropic's own engineering guidance on effective context engineering for AI agents reframes the agent context as a scarce, perishable resource—one that must be curated, compressed, and explicitly garbage-collected.
The Agent Memory Hierarchy: Five Tiers, Modeled on CPU Caches
Treat the context window like an L1 cache: small, fast, expensive, and backed by progressively larger, slower tiers. Production-grade coding agents implicitly implement this hierarchy.
Tier 0 — Instruction Memory: System Prompts and AGENTS.md
Static, always-resident directives: build commands, test invocation, code style, architectural boundaries. This tier should be deterministic and version-controlled—never generated per-session, because any change invalidates the prompt cache and forces full-price reprocessing of the entire context.
Tier 1 — Working Context: The Active Token Budget
Currently open files, recent tool outputs, and the running task plan. Discipline here is brutal: load file slices (function-level, not file-level), truncate stack traces to the deepest application frames, and strip node_modules-style noise before it enters the loop. A working set exceeding ~30–40% of the window is a reliability risk, not an asset.
Tier 2 — Retrieval: Grep-First, Embeddings-Second
Counterintuitively, agentic grep/glob search outperforms embedding-based RAG for code repositories in most agentic benchmarks. Code identifiers are high-entropy, discrete tokens—ideal for exact-match retrieval and hostile to dense-vector similarity. Hybrid architectures win: embeddings for conceptual queries ("where do we handle auth failures"), ripgrep-style tools for structural queries ("every caller of refreshToken").
Tier 3 — Durable Project Memory
Persistent notes that survive across sessions: architectural decision records, known pitfalls, the rationale behind non-obvious code. The key design constraint is freshness enforcement—stale memory that contradicts the current codebase actively corrupts agent behavior. Treat memory files like a cache: validated against reality, or evicted.
Tier 4 — Sub-Agent Contexts: Isolation as a Feature
Spawning isolated sub-agents gives each a clean, empty L1. The orchestrator dispatches a narrow task; the sub-agent burns its own tokens on noisy exploration; only a distilled artifact returns.
Compaction: Garbage Collection for the Context Window
When working context nears budget, agents must compress history rather than truncate it. Naive summarization fails because it discards the exact identifiers the agent still needs. A robust compaction schema must preserve:
- Task intent and acceptance criteria — the original goal, restated verbatim.
- Decision ledger — every architectural choice with one-line rationale (prevents agents from re-litigating settled decisions).
- Exact symbol inventory — file paths, function names, and line numbers already created or modified.
- Open TODOs and known-broken state — failing tests, incomplete refactors, unverified assumptions.
Hook-driven compaction (triggered at a token threshold, executed by a dedicated summarizer model) outperforms single-model auto-compact because the summarizer can be a cheaper model running against a structured extraction prompt rather than free-form recall.
Sub-Agent Delegation Patterns That Actually Reduce Token Burn
Delegation is not parallelism theater; it is context quarantine. Three patterns dominate production systems:
- Orchestrator–worker: The main loop holds only plans and distilled results. Workers handle file exploration, log analysis, or test triage and return a verdict object, never raw output.
- Read-only scouts: A sub-agent greps, reads, and maps the relevant code surface, returning a curated briefing. The orchestrator never sees the thousands of tokens the scout discarded.
- Verification isolation: Test execution and CI log parsing run in a sandboxed sub-agent; only pass/fail plus minimized failing assertions re-enter the primary context. This single pattern routinely cuts orchestrator context growth by 50%+ on complex tasks.
AGENTS.md: The Emerging Standard for Machine-Readable Repos
Repo-level instruction files have consolidated around the AGENTS.md specification, now supported across major coding agents and IDEs. Effective files share measurable traits:
- Under ~500 lines—every instruction token is a recurring per-request tax.
- Commands, not prose: exact build/test/lint invocations the agent can execute verbatim.
- Explicit boundaries: "never modify generated/", "always run pnpm typecheck before declaring done."
- Placement at repo root and in subdirectories, where nested files override parents—mirroring .gitignore semantics.
Measuring Context Efficiency: Metrics That Matter
Optimizing without instrumentation is guesswork. Instrument these per-task metrics in any agentic pipeline:
- Tokens per resolved issue — the core efficiency ratio; track cache-hit percentage separately, since cached input can cost ~10% of uncached.
- Context saturation curve — token count plotted against task progress; a sawtooth (growth → compaction) is healthy, monotonic growth predicts failure.
- Instruction adherence rate — fraction of completed tasks where AGENTS.md constraints were honored; decay signals compaction dropping Tier 0 facts.
- Tool-call depth per sub-task — a proxy for whether retrieval (Tier 2) is doing its job or leaking into the working set.
Implementation Checklist for Production Agents
- Version-control AGENTS.md; never mutate system-prompt tokens mid-session.
- Cap working context at ~35% of the window; trigger compaction at threshold, not at overflow.
- Default to grep-based retrieval; reserve embeddings for conceptual queries.
- Route all raw tool output through sub-agents; return structured artifacts only.
- Maintain a decision ledger and exact symbol inventory across every compaction cycle.
- Log the full context-efficiency metric set per task and gate deploys on adherence rates.
Context engineering is now a first-class architectural concern—on par with database schema or API design—for any team shipping agentic software. If you're building LLM-powered products and want a partner fluent in these patterns, explore our AI-native development services, or go deeper on applied agent architecture in our engineering blog.