Why Your Agent's Biggest Bill Is Invisible
Every agent has a mechanism that decides what its model sees on each turn. In most production systems, that decision was frozen during prototyping and never revisited. That is a problem, because it usually drives the largest share of operating cost — and quietly contributes to disappointing answers.
Here is the subtle part: this is the one piece of an agent that can improve on its own. The model stays as capable as the day you picked it. Instructions only change when a human rewrites them. But what an agent knows, can access, and remembers grows as it runs. Managing that growth is called context engineering.
A model has no memory of its own. On every turn, the context window hands it everything it can use: instructions, tools, retrieved documents, conversation history. When the turn ends, that context vanishes and must be rebuilt for the next one. For a chatbot answering one question, this is fine. For an agent working across dozens of turns toward a single outcome, it is often the single largest expense — and unnecessary content gets billed repeatedly.
The less obvious cost is quality. More context does not mean better answers. A relevant fact buried in 40 pages is harder to use, and a bloated tool list makes the wrong choice more likely. Each mistake adds turns — and cost — to recover. Unlike swapping to a cheaper model (which trades away quality) or shortening instructions (which can weaken answers), removing unnecessary context lowers cost without lowering quality. That makes it the easiest optimization a team can actually defend to leadership.
(Source context: the economics of agent optimization)

The Four Questions That Define Context Engineering
Context engineering is not a one-time design decision — practiced continuously, it is how an agent improves, because every turn reveals what it actually used. Four questions cover the work, and most teams tackle them in this order.
1. What should the agent know?
Many teams start with broad search that dumps whole documents into the prompt. Easy to build, expensive to run. A managed knowledge layer fixes this by decomposing a query into subqueries, searching sources in parallel, semantically reranking results, and returning only grounded passages with citations. In internal benchmarks on BrowseComp-Plus, this approach improved evidence recall by up to 54% while cutting retrieval token costs by 34%.
2. What should the agent be able to reach?
Tool overhead is the sneaky one. Adding a tool is one line of code, but its full description rides in the prompt on every turn — needed or not. When a toolbox grows to hundreds of tools, that description tax dominates your bill. The fix is tool search: instead of exposing the full list, the model gets a way to describe what it needs in plain language, and a way to call whatever comes back. The cost of the tool list stays flat, however large the toolbox grows.
# Naive: every tool description is sent on every turn
tools = [search_tool, code_tool, sql_tool, ...] # 200+ tools
response = model.chat(prompt, tools=tools) # 💸 full description tax every call
# Better: expose a single tool-search entrypoint
tool_search = {
"name": "find_tool",
"description": "Describe the capability you need in plain language."
}
response = model.chat(prompt, tools=[tool_search])
# The model asks for a capability, you resolve it server-side,
# and only the matched tool's schema enters the next turn.
Internal benchmarking against a public open-source tool-retrieval dataset showed average input-token consumption dropping roughly 97% for large tool libraries — a direct hit to inference cost.
3. How should the agent do the work?
Knowledge and tools cover what an agent can find and do. Neither covers how your company expects the work to be done — the escalation path a support agent follows, the checklist a code review applies. That guidance usually lives in instructions, gets copy-pasted across agents, and ships on every request even when irrelevant. The fix is a skill: a named, centrally managed procedure that an agent loads only when relevant. It sees the skill's name and a short description first; the full instructions load on demand.

What the agent should remember — and the limits of all this
Agents need continuity, but they do not need to carry every detail from every interaction. Replaying an entire conversation to the model adds cost and burns context, even when only a few details still matter. A practical memory model splits into three tiers:
- Session memory — the current conversation.
- User memory — preferences and facts that persist across sessions.
- Procedural memory — learned workflows and task execution patterns.
Procedural memory complements centrally managed skills nicely: a skill defines the approved procedure, while procedural memory lets an agent learn from its own execution. In one set of internal evaluations, enabling procedural memory produced roughly a 5% improvement on STATE-Bench and Tau-Bench.
⚠️ Where this gets hard
Context engineering is not free lunch. A few honest caveats:
- You need observability first. You cannot prune what you cannot see. If you do not log what enters the context window each turn, you are guessing.
- Retrieval quality is a ceiling. If your knowledge base is stale or your reranker is weak, no amount of pruning will save answer quality.
- Memory is a compliance surface. User memory and procedural memory store data. Retention policies, TTL, and user-level isolation are not optional — they are the price of admission in regulated industries.
- Vendor lock-in is real. Managed knowledge layers, toolboxes, and memory services are convenient, but design your agent so retrieval, tools, and memory are swappable behind interfaces. Frameworks like LangGraph, the Microsoft Agent Framework, GitHub Copilot SDK, and Claude Agent SDK are all reasonable targets — do not couple your business logic to one vendor's SDK.
- Benchmarks are directional, not promises. A 97% token reduction on a public tool-retrieval dataset is a signal, not a guarantee. Measure on your workload.
Next steps to learn
- Profile a single production agent: log the token breakdown of instructions, tools, retrieved docs, and history per turn. The biggest slice is your first target.
- Replace static tool lists with tool-search for any agent with more than ~20 tools.
- Move repeated procedures out of instructions and into versioned, centrally managed skills.
- Add session/user memory only after you have retention and TTL policies in place.
- Read up on agentic retrieval and semantic reranking — they are the two biggest levers on the knowledge side.

The Bottom Line
Context engineering is not about shrinking prompts. It is about building agents that get better with use. When the knowledge they draw from, the tools they discover, the procedures they follow, and the memories they retain all improve over time, an agent becomes both more capable and more efficient — without a rebuild.
The practical takeaway for any team shipping agents today: before you swap models or rewrite prompts, look at what enters the context window on every turn. In most production systems, that is where the money is going.