← Shu Yang

Two New Concepts for Frontier AI Systems: “Context Engineering” and “Chain of Command”

Originally published on Notion, where the referenced figures can be viewed.

Introduction

The key objective of AI systems, whether large language models (LLMs), multimodal models (MLLMs), or multi-agent systems, is to ensure that these systems maximize helpfulness and freedom for builders, developers, and users, minimize harm, and choose sensible defaults [3]. As AI systems grow increasingly complex, the external information they must process from diverse environments, roles, and instructions also becomes more intricate and may even conflict. This raises two fundamental questions:

To further investigate and optimize these challenges, two emerging concepts have been proposed: Context Engineering and Chain of Command. In this blog, I introduce their definitions and explain how they differ from existing related concepts in the field.

Context Engineering

Agents need context to perform tasks. Context engineering is the art and science of filling the context window with just the right information at each step of an agent’s trajectory. [1]

It refers to the set of strategies for curating and maintaining the optimal set of tokens (information) during LLM inference, including all the other information that may land there outside of the prompts. [2]

The difference between prompt engineering and context engineering

Prompt engineering: methods for writing and organizing LLM instructions for optimal outcomes. The primary focus is how to write effective prompts, particularly system prompts.

Context engineering: as we move towards engineering more capable agents that operate over multiple turns of inference and longer time horizons, we need strategies for managing the entire context state (system instructions, tools, Model Context Protocol (MCP), external data, message history, etc.).

Why we need context engineering

A vivid framing from LangChain: LLMs are like a new kind of operating system. The LLM is like the CPU and its context window is like the RAM, serving as the model’s working memory. Just like RAM, the LLM context window has limited capacity to handle various sources of context. And just as an operating system curates what fits into a CPU’s RAM, “context engineering” plays a similar role.

Complex agents likely get context from many sources: the developer of the application, the user, previous interactions, tool calls, or other external data. Pulling these all together involves a complex and dynamic system. At one end of the spectrum we see brittle if-else hardcoded prompts, and at the other end we see prompts that are overly general or falsely assume shared context. We need context engineering to make sure the right information, tools, and instructions are used for inference.

How can we do context engineering?

Following the taxonomy in [1]:

The chain of command

Is merely ensuring the involvement of the “proper” context in the model’s or agent’s inference (which we can never fully guarantee anyway) sufficient to achieve the objective mentioned in the introduction? No. The model or agent also needs to know how to use this context and how to resolve potential conflicts, which means it needs to know which instructions, information, and sources have higher privilege. This is also an important and foundational aspect of model safety and alignment.

Chain of command is the structured principle that determines how the model should prioritize and reconcile multiple or conflicting instructions in a coherent and consistent manner.

Instructions and levels of authority

Taking the OpenAI Model Spec as an example, instructions are organized into levels of authority:

How assistants are trained to follow this authority level

Why the chain of command is important

“Today’s LLMs are susceptible to prompt injections, jailbreaks, and other attacks that allow adversaries to overwrite a model’s original instructions with their own malicious prompts. In this work, we argue that one of the primary vulnerabilities underlying these attacks is that LLMs often consider system prompts (e.g., text from an application developer) to be the same priority as text from untrusted users and third parties.” [4]

If the model follows the privilege levels above, it is far less likely to be hijacked by, say, an injected instruction in a web search result. A chain of command helps the model safeguard itself against threats such as (see Appendix B of [4] for examples and related datasets):

It also helps address the specific risks enumerated in the OpenAI Model Spec:

Resources

  1. Context Engineering for Agents. LangChain blog. blog.langchain.com/context-engineering-for-agents
  2. Effective context engineering for AI agents. Anthropic. anthropic.com/engineering/effective-context-engineering-for-ai-agents
  3. OpenAI Model Spec (2025-09-12), Chain of command. model-spec.openai.com
  4. Wallace, Eric, et al. “The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions.” arXiv:2404.13208 (2024). arxiv.org/abs/2404.13208