Prompting

What is context engineering? vs prompt engineering

A model can only work with what is in front of it. That sentence is the whole idea. Context engineering is the work of deciding what is in front of it, in which order and in what shape.

Prompt engineering vs context engineering#

Prompt engineeringContext engineering
ScopeOne requestEvery request in a system
QuestionHow do I word this?What should the model see right now?
LeversInstructions, examples, formatInstructions, retrieved data, tools, memory, history
Typical failureVague or conflicting instructionsWrong, missing or bloated information
Who does itWhoever writes the promptThe system, every turn

They are not rivals. A clear prompt is one part of the context.

What is in the context window#

System instructionsrole, rules, format (stable, cacheable)Tool definitionsnames, descriptions, input schemasRetrieved knowledgedocuments, search results, filesMemoryfacts carried across sessionsConversation historyearlier turns and tool resultsCurrent user messagewhat is being asked now
Every request assembles these layers. You control what goes in each one.

Why it matters most for agents#

A chatbot gets one question. An agent runs for dozens of steps, and every tool result lands in the context. Left alone, it fills with stale search results, huge file dumps and old reasoning. The model gets slower, costs more and starts missing the one fact that matters.

Typical agent failures are context failures:

  • It lacked a fact (so it guessed).
  • It had the fact buried under 30 irrelevant tool results.
  • Two instructions contradicted each other.
  • A tool description was vague so it picked the wrong tool.

Seven practical moves#

  1. Retrieve, do not dump. Fetch the three relevant chunks, not the whole wiki. See what is RAG.
  2. Write tool descriptions like docs for a stranger. Say what it does, when to use it and what it returns. The description is how the model chooses.
  3. Trim tool output. Return the fields needed, paginate, truncate with a note.
  4. Summarize or clear old turns. Long histories need compaction or clearing.
  5. Put stable content first. Instructions and tool lists at the front can be cached; volatile data goes last. See prompt caching.
  6. Separate data from instructions. Mark retrieved text clearly (for example in XML tags) so the model treats it as material, not commands.
  7. Write the always-true facts down once. In a coding agent that is a CLAUDE.md.

A quick example#

Bad context for a support bot: the entire 80-page policy PDF plus the last 40 messages.

Better: the system prompt, the three policy paragraphs retrieved for this question, the last four messages, and the customer's plan tier from a tool call. Smaller, cheaper and more accurate.

How to debug it#

When an answer is wrong, do not rewrite the prompt first. Print the exact context the model received and ask: could a smart human, given only this, answer correctly? If not, the fix is in the context, not the wording. This is also what evals are for.

Frequently asked questions

What is context engineering?

It is the practice of choosing and structuring everything an LLM sees at inference time: the system prompt, retrieved documents, tool definitions, conversation history and memory, so the model has the right information in the right format.

Is context engineering replacing prompt engineering?

It is a broader view of the same job. Wording your prompt clearly still matters, but in apps and agents the bigger lever is what you put in the context window and what you leave out.

What is a context window?

The maximum amount of text, measured in tokens, a model can consider in one request, covering the input and the output.

How do I reduce context problems in an agent?

Retrieve only what is needed, trim old tool results, summarize long histories, keep stable instructions at the front so they can be cached, and write short, specific tool descriptions.