Concept 09 · 7 min
Anything it reads can give it orders.
Scroll to start
Scene 1
Data or orders?
Sarah’s assistant can read her emails, the CRM and account statements, and it can send emails. She asks it to summarize Mrs. Chen’s emails.
[Content TODO]
Scene 1: the example email as Sarah sees it, the prediction, and the reveal (see specs/09-limits.md, Scene 1).
To an LLM, everything in its context window is text: Sarah’s request and a stranger’s email look alike. Text that slips orders into what it reads is called a prompt injection.
Scene 2
Same email, one brake
Same email, same request. One change: the program now asks Sarah before anything leaves the bank. You are Sarah.
[Content TODO]
Scene 2: the step-by-step run with the program's fence and the reader's decision, Deny or Approve (see specs/09-limits.md, Scene 2).
That brake is called human approval. It wasn’t in the LLM’s instructions; it was in the program: nothing leaves the bank without Sarah.
Scene 3
Three abilities, one risk
A leak needs three things at once. Switch each ability on and off.
Switch the three abilities on and off, and watch when a leak becomes possible.
One side is missing: nothing can both reach the data and carry it out.
Private data, outside content, a way to send things out: together they make the lethal trifecta. Each guardrail cuts one side; giving only the access a task needs is called least privilege.
Scene 4
What you never paste
Five things Sarah might paste into an assistant. Where can each one go?
Drag each item into the right column: “Any assistant”, “Only the approved one” or “Never, anywhere”.0/5 sorted
No mouse? Tap an item, then a column.
Before pasting, ask: would I hand this to a gifted stranger? Your company’s policy says which tools count as approved.
Hype vs reality
- What the hype says
“Just tell the AI ‘never do X’ and it won’t.”
What actually happensInstructions in the prompt can be overridden by text it reads. Real limits live outside the model: permissions, allow-lists, approvals.
- What the hype says
“Each tool is safe, so connecting them all is safe.”
What actually happensThe danger is the combination: private data, outside content and a way out. Before connecting an assistant, check whether it has all three.
- What the hype says
“A smarter model will be immune to tricks.”
What actually happensSmarter models resist better, but anything that reads untrusted text can be steered. Design as if it will be.
Under the hood
The checks that matter live in the program.
step = agent_step(prompt, tools) # only the tools the task needsif step.sends_email: if step.recipient not in KNOWN_ADDRESSES: # allow-list: checks the address step = blocked("unknown recipient") elif not sarah_approves(step): # human approval: checks the intent step = blocked("Sarah said no")Prompt injection is the first risk on OWASP’s reference list for LLM applications (OWASP Top 10 for LLM Applications). The lethal trifecta is Simon Willison’s name for the combination (“The lethal trifecta for AI agents”).
In 3 sentences
- To an LLM, everything it reads is text, so an email or a file can slip in orders.
- Never give an assistant private data, outside content and a way out at once without a brake; the safest brakes live outside the model.
- Never paste what you wouldn’t hand a gifted stranger, and check what matters.
Did this make sense?
0/9 concepts unmagicked
Go to the challenge- 01LLMAvailable · How can it write so well without understanding? · Available
- 02ContextAvailable · Why does it forget what I said earlier? · Available
- 03PromptAvailable · Why are my results mediocre? · Available
- 04HallucinationsAvailable · Why does it make things up? When can I trust it? · Available
- 05RAGAvailable · How do I make it use our internal documents? · Available
- 06ToolsAvailable · How can it act, not just talk? · Available
- 07AgentAvailable · What is an agent, concretely? · Available
- 08SkillsAvailable · How do we specialize it without retraining? · Available
- 09Limits & safetyAvailable · What must I never trust it with? · Available
Solve the sandbox challenge to unmagic this concept.
End of the journey
Three habits to keep
- 1Check what mattersNumbers, names, references, and the sources it read.
- 2Give the least access neededTools that look by default; never all three sides of the trifecta without a brake.
- 3Keep a human before what can’t be undoneSending, paying, deleting. Urgent requests that look internal deserve a second look.
Get notified when it’s out:
Scripted scenario, not a live model.