Get started
AI engineeringContext engineeringAI agentsMemoryEvidence

Context engineering is not prompt engineering

Adam Cowles2026-10-03T18:30:00.000Z5 min read
Layered translucent planes of documents, memory and tool output converging into a single witnessed working set, cyan and violet over near-black

A prompt is what you say to the model. Context is everything the model gets to see when it answers. Almost all the hard work has moved to the second one.

A prompt is a sentence. Context is a working set: the documents, the memory, the earlier steps, the tool output, the few examples, all of it assembled and handed to the model at the moment it has to decide something. Prompt engineering was the craft of wording one instruction well. Context engineering is the craft of building that working set, and for an agent that takes actions rather than just answering questions, it is now where most of the quality lives.

The shift is not fashion. It followed the context window. When a model could only see a couple of thousand tokens, phrasing was the lever you had. Now that a model can see hundreds of pages, the lever is selection: what goes in, in what order, and what you deliberately leave out. Give an agent too little and it guesses. Give it too much and it drowns, loses the important line in the middle of the haystack, and costs you more per call to do it. The skill is curation, retrieval and ordering, not wording.

What the working set is actually made of

Pull apart the context behind a single agent step and you find layers, each from a different place and each with its own failure mode.

  • The instruction: the system prompt and the task. Stable, small, still worth getting right.
  • Retrieved knowledge: the documents or rows fetched for this specific task, usually by a similarity search. Wrong retrieval here is the most common cause of a confidently wrong answer.
  • Memory: what the agent or the user established earlier, in this session and in previous ones. Skip it and the agent is an amnesiac that re-asks what it was told yesterday.
  • Trajectory: the steps the agent has already taken this run, and what each returned. This is what keeps a multi-step task coherent.
  • Tool output: the live results it called for, which can be large and need trimming before they go back in.

Context engineering is the set of decisions that assembles those layers for one step, under a token budget, so the model sees enough and not too much. None of it is visible in the prompt. All of it determines the answer.

A worked example

An agent is resolving a support ticket: "my export is empty since the weekend." A weak context set hands the model the ticket text and the whole product manual, and hopes. A well-engineered one assembles something closer to this: the ticket, the customer's plan and their last three tickets (memory), the two manual sections that similarity search tied to "export" (retrieval, scoped), the result of a live status check on the export job (tool output, trimmed to the failure line), and the fact that the agent already tried a cache flush and it did nothing (trajectory).

Same model, same prompt. The second agent answers from the customer's actual situation; the first pattern-matches the manual and suggests a cache flush that was already tried. The difference is entirely in the context, and building the second set well is the engineering.

The part almost nobody engineers: proving what the context was

Here is the question that arrives later, when the agent got it wrong and moved real data or approved a real payment: what did it actually see when it decided that?

In most stacks there is no answer. The working set was assembled in memory, passed to the model, used, and discarded. The retrieval that pulled the wrong document left no record of which document. The memory that was stale when it mattered is already overwritten. You can read the final output and the original prompt, and reconstruct nothing in between. For a chatbot that is tolerable. For an agent that acts, it means the most important audit question has no answer, and "the AI did it" is the end of the investigation rather than the start.

This is the honest gap in the whole context-engineering conversation. The field is getting good at building the working set. It has barely started on being able to show, after the fact, what the working set was.

Context you can reconstruct

The fix is not another framework. It is treating the pieces of the working set as events worth recording, not as scratch memory. If every memory read, every retrieval and every tool call is a witnessed event with a timestamp and a hash, then the context behind any decision is reconstructable from the record instead of lost with the process that held it.

That is the line Quox builds on. Brain2 is the governed memory and knowledge layer an agent reads its context from, and Quox records agent actions as tamper-evident events rather than log lines you have to trust. The point is not that memory makes an agent smarter, though it does. The point is that when the context is governed, the question "what did it know when it decided that" has an answer you can hand to an auditor, because the reads that built the context are on the same record as the action that followed. That is the same idea as a cryptographic audit trail for AI agents, pushed one layer earlier: not just what the agent did, but what it was looking at when it chose to.

Context engineering is already table stakes. Any team shipping agents is learning to build the working set. The next bar, the one that matters once an agent is allowed near anything that costs money or breaks trust, is context you can prove. You can see how Quox makes agent decisions checkable on the evidence layer, and the wider agent picture on the agentic AI hub.